Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:54.027500Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 100 of 167 outbound references and 0 inbound Pith citation observations for arXiv:2608.03092.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:54.027500Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 167 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b44c3712-e4fb-4d25-b042-b058b40d0c96 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Approximating
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d978763-eb09-4404-ac3a-ea6d751c60d3 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Thinking Machines Lab: Connectionism , year=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bc3d340-d0ad-46f6-886e-05ad422f23b4 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f930e07-0f31-4a88-86bd-77b11984480b · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b9bebc-e6c2-4376-a935-351d30f38d02 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b0f2f6-ed7b-432d-b120-1e6766403ba6 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be70df40-b56d-4ea7-a077-17a1b2023331 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0787fa9-408b-4eb8-9531-1affac13a384 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31bca12-0bbb-4470-bd94-c9265e587d90 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation NIPS Deep Learning and Representation Learning Workshop , year=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c332071-41cd-4d27-a618-1b40ac4a1c61 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66007b7a-c580-4046-af05-70c4019a05ec · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e772e42-ca45-4703-972f-f65ccf585771 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fb0731-3699-499a-973e-7019f2ddc652 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301e8b44-978d-47bb-8ef6-e89e0d9de9b6 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea2d3af1-68aa-4862-81bf-184310a1320e · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bffd4d0-e002-4b5a-8291-bc548fcdbf54 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e160db9-9907-4243-9747-5cd4c6284f54 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9981f875-77c8-4e0f-a65f-c41b1036962d · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Language Models are Super
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8609371-1e60-447a-afec-bf88a774d8f6 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c3a387-c06f-4e35-98e1-48f32b788dae · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 39th International Conference on Machine Learning (ICML) , year=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb26a02-3136-4743-be38-7b335f5a47f4 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Constitutional AI: Harmlessness from AI Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55ff134-a8c9-4fef-be53-5822847b22cf · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f469fe5c-7c2f-4362-a77f-f9cb0ec5a8f0 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a04f551-0c78-48aa-bcd0-2e3b1650c66a · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Hashimoto , year=
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5ba99a-6b97-4694-a945-a9535aee84ac · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Arithmetic Control of
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c6e0cdb-3ef0-409f-b06d-7d3837124cd8 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Rewarded soups: towards
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5911d5ed-af19-4c20-99d2-f54f52ae8c3a · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55315f55-352e-45fb-92d4-c8072d293dd1 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Back to Basics: Revisiting
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e845bf90-36b6-4458-8001-5b3997453d2c · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f21dcfbf-6f20-4b6c-b3b0-a49a4d561e47 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2606.16771 , archivePrefix=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a56a4a33-f71d-4464-8afd-e78f8f8071ad · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 947a9e8c-6c31-4e97-b98e-94b3b52ea8b0 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56257440-42c7-4d8e-8768-c59fd0c926f5 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60362f9e-c83b-4f2e-b4c4-f50e93c95758 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1355564d-ac98-4810-a57f-0b906cb5eace · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Hao Zhang and Ion Stoica , booktitle=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dab2f60-a705-448b-a87e-589b931b20b9 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3971191-ef85-4ff8-b1c6-f11cee07edff · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Patil and Tianjun Zhang and Xin Wang and Joseph E
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7f4ffe-1f1a-4a0c-9444-fbd9de0231a0 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d237aab3-d2f0-45e2-8d98-a83959f16682 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c73846-1a10-4f50-9319-e3fb17d0e450 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5f1707d8-0451-4278-be2d-ea5f5f65f148 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7100e852-0073-42bc-9aef-0a21d22ff9ef · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458571d0-4753-49b3-bc5e-b90ebef85fe3 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b74d5406-0524-46d2-8a75-34d0a3b1ce8c · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: ACL 2024 , year=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d38c34-105f-4271-84ec-e8728d765d87 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Singh and DJ Strouse and Tuomas Sandholm and Ruslan Salakhutdinov and Anca D
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a920b9b-be2f-48cc-8f55-8369bda27d7d · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b17cf8-95dd-4af2-a4f0-cf5d12562fa7 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , doi=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910b124d-5a82-4792-b21f-92dba7479cf3 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66bf340-d024-4cef-9b55-5130159e6940 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Measuring Mathematical Problem Solving With the
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d649c44-756f-4c84-8759-1a43d032246e · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 739164c7-0b3f-434f-b1dd-9d7647fb01c4 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933ed797-10b5-4678-bba9-08d2b4935b00 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Machine Learning , pages=
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c6525fa-6a6e-43e7-b798-c2b3bd17038b · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53692ec1-9980-46a6-ac36-d69d18c923d9 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc3c0ad2-cb95-434c-b0b0-3abcff808756 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1506ac23-0a71-46b5-a02c-d10d8f928cfd · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d606a231-f2ae-4509-8597-e32f247ea1ee · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Ion Stoica , booktitle=
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2370deaa-773b-479e-b087-935cb248bd00 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Machine Learning , volume=
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4003c66c-0810-457c-b427-a3be5aa1dc7d · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96673f63-9191-467a-80d1-a86c82d45f75 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the association for computational linguistics: EMNLP 2024 , pages=
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb436b5f-bffc-44f3-8219-b28e115e373f · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Everyone Deserves A Reward: Learning Customized Human Preferences
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f84292-f9a6-42c0-938a-43315dd4d88a · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Kimi-VL Technical Report
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c547009-05c7-4ae9-8842-3864d34f42c6 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d3048e-1d8c-4e85-be5b-7b2308f1e963 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c9e7be-5e37-4de4-a6ae-018e2e77da91 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2000 , eprint=
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed62110f-f4e8-4f0a-8f43-9a5398d88936 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 1997 , publisher=
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a830bc8-5b09-4916-afbd-fc7199199291 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74fe22e8-2260-483f-be44-6ee4dc7feb15 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff4ac2d3-d624-4216-bc39-52052533cbe1 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Representation Learning with Contrastive Predictive Coding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8595768-a948-4f78-a4aa-69e7e79a006a · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b0a12de-5c06-49c9-b42c-d1066999ad54 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efceb99c-e966-482d-88be-1bee64ba9100 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1911532-51fd-466a-9c76-2074d31e8c6b · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation A Long Way to Go: Investigating Length Correlations in RLHF
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09f95b76-65fa-442f-af2a-decc09c41a21 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Towards Understanding Sycophancy in Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f904c68f-6879-4859-8ce7-4640b3670e3c · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef0411b-074d-420d-b9bb-4b3e3314c8c2 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90036455-6a28-4aa9-8e78-d0208b1910d2 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdda4da-491a-432d-a3a1-4b8f546c4c3d · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8724f47-7509-4199-9099-2e2d18c147e5 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 783e67f2-7ba2-4956-96bd-cd9d5f43d6f6 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283518ca-0f74-43ac-9f73-68e4f8110187 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GPT-4o System Card
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c8126d-f837-483d-947e-fbd78482a8dd · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a4ea9b-7e5b-4972-8c5a-333a96a32c0f · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2023 , eprint=
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f68f78d-087d-498d-8e1d-b1b9609f78ce · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd26a059-5d3c-4d27-828a-4c16ff5ba1a5 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c52bc28-2c74-4fa0-9f45-d1f6cf46fb42 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d7d7669-d088-4acb-b031-e45a426084a8 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d1257dc-eee2-430c-9915-3f920322fb5b · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , booktitle=
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98600c1f-5a41-4af8-ae8b-0c0c8d4267af · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , booktitle=
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaeff446-e299-471c-bef7-f32774ea6311 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9157bab-d55e-46d8-972f-18644528a140 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The method of paired comparisons , author=
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3671c0d6-78d4-49cf-9a2f-e02ec65f551a · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db5d0e8-e58a-437c-85d2-fc3a6e5e6409 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3c386e-3536-4b72-9788-809f86b0f89b · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6661f9bc-b9bc-4d85-a316-01b58d24a03f · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c4114e-0800-4c37-9b5e-fc611cca1038 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Instruction Tuning with GPT-4
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4702e3-bdb3-4135-b817-af21f4a2d47a · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1884369b-ba39-47b9-86d7-cc2fd09e02fb · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a58c936d-b789-4294-8d0b-8761a1b6fc64 · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d8a1c50-2af2-488e-b1cd-26f9a98fee5a · outbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.