Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:57:59.957126Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.29246.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:57:59.957126Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2252a65d-8797-4a2d-88ac-2bce9c40056d · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c56373e8-8d14-4135-aa16-9e83e4560f72 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c428343-9b71-46ba-97f9-9d108374649c · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Constrained policy optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 710a68b9-ed44-4579-a283-d54ab2267edc · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL A General Language Assistant as a Laboratory for Alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c886ddbb-9dfd-40bf-878c-c32ef9e2f118 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f536a981-51e5-491c-8815-21f3cb3dd3b8 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Safe rlhf: Safe reinforcement learning from human feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00b266f5-8f52-4085-974c-087954635c6e · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dynamic multi-reward weighting for multi-style controllable generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b1c5c6-eaf4-459d-b99d-87d16faeede5 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Controlled Text Generation via Language Model Arithmetic
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b35f7ae-9383-43d2-871d-d0bd26217ed8 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Sciknoweval: Evaluating multi-level scientific knowledge of large language models.arXiv preprint arXiv:2406.09098, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76e35d0d-e00e-4726-a1b9-7acb6c8fad02 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aeeb0ac-8a3e-44c7-b16d-220ba9594999 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Training products of experts by minimizing contrastive divergence.Neural computation, 14(8): 1771–1800, 2002
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d09ef992-9976-4b7a-ad03-d54637359612 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283c1b82-022e-4dba-9eee-eede598c0b17 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0ac92d-9ac7-47ee-aece-ac1a95b36dc2 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL The power of scale for parameter-efficient prompt tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f23978-ec55-455c-a541-bddc6d2be2bc · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Gradient-adaptive policy optimization: Towards multi-objective alignment of large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2338c79f-231e-4e2f-87db-97c88a8d3dd1 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Prefix-tuning: Optimizing continuous prompts for generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c44bc475-1f86-4cf5-814a-33156b0f5d9a · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dichotomous Diffusion Policy Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4338bba-2e77-45b0-bc39-7ddafa50bf5d · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Mitigating the alignment tax of rlhf
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43580a14-00a4-4046-a1fe-0e31a0e10efd · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL DeepSeek-V3 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f00b3543-9e56-4633-99c3-50c5359b8567 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dexperts: Decoding-time controlled text generation with experts and anti-experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e8983d-011f-4b77-9d09-18d24677f944 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d09a2950-006e-4929-a5c3-06a546bc80b5 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Contrastive Decoding Improves Reasoning in Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 519e1b21-465a-48ff-8af6-285222c09a1f · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79d03c10-79d7-4257-9e56-85f8bade94ba · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60fae710-b15f-47ad-aaab-3b0a667bc766 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Efficiently Scaling Transformer Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4dbf456-c303-4424-8266-63279c71fa1d · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Toolrl: Reward is all tool learning needs.Advances in Neural Information Processing Systems, 38:105523–105553, 2026
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a6c740-2c46-41a6-bc98-84658d90fbc1 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dmoerm: Recipes of mixture-of-experts for effective reward modeling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933e819f-9aaa-4f02-9e7d-0e57b9f5a0f9 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Direct preference optimization: Your language model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d164ed71-f4cc-46b1-90dc-f45e60b3ff22 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL WARM: On the Benefits of Weight Averaged Reward Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddb18dcd-5389-493e-9ace-2059d88a841b · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8628402e-3a1b-475d-9b05-e8d74a491354 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL A survey of multi-objective sequential decision-making.Journal of Artificial Intelligence Research, 48:67–113, 2013
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 357c0212-c5f2-4562-9d4a-08842e2ba2c8 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries, 23(3):289–301, 2022
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf1a97c0-b3a4-4c86-a02c-b6ddc315f087 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Proximal Policy Optimization Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8885a5af-f4aa-4145-b41c-ea459b4ef695 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c2228c-b2e3-4d21-969e-27d1be6177fe · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL MIT press Cambridge, 1998
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9549157-ac81-43e7-9ea4-cae6e2590bed · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Hashimoto
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 869403b3-bd32-4e28-859a-bedd41bbcac0 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Qwen2.5: A party of foundation models, September 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15785910-7208-433a-9312-1489c8754843 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Interpretable preferences via multi- objective reward modeling and mixture-of-experts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2059eb48-5b8f-409d-95e3-8eda809e0613 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Aligning Large Language Models with Human: A Survey
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53e6144e-33e8-454b-b059-9793794acb39 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Multi-objective Reinforcement learning from AI Feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f1d597-9f39-42f4-9f2b-210396983f6f · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688eef7b-a850-4b9d-9991-2ddcb722e34f · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f28c77-d3a1-45e7-a0ca-68f330b176c2 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Dapo: An open-source llm reinforcement learning system at scale.Advances in Neural Information Processing Systems, 38:113222–113244, 2026
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 236c7969-d734-4214-84c0-fce3c1900fdd · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c432441f-4f02-4915-8c54-66fd1168a07a · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Group Sequence Policy Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd437aff-d4f8-4944-b63b-fbeec4b8b18c · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Beyond one-preference- fits-all alignment: Multi-objective direct preference optimization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a908d3b-3316-4cfe-92db-d48ff77ec12e · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Fine-Tuning Language Models from Human Preferences
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f791c3-5068-46c1-b111-0ee663959998 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e665d08-12e3-4708-b27b-49ad6f654631 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edc1baf-0eab-471f-a361-e479bf5a0652 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0baae10-9d9a-48d0-a9b6-0f1a15126da5 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40321411-77e6-48c6-ae9f-75596065d0c4 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL A", "B",
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce02a277-1602-412e-abba-2d5046ed19cf · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Provide at least one of<tool_call> or <response>
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a755a497-cc37-4d2a-88fe-99d37282fd9d · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL name" field and a
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe4842b0-bf39-4b9c-b21d-b92771163819 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a64dbc0-e2c5-4ab3-911e-4f6f06e01b83 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8af99f-cb91-4ad2-8e21-6ff0c6274339 · outbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL name ":
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.