Pith. sign in

Paper Citation Record · LEDGER

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

As of 16 August 2026, this Paper Citation Record lists 100 of 167 outbound references and 0 inbound Pith citation observations for arXiv:2608.03092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03092 v1

Coverage vector

measured 100 of 167 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:54.027500Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 167 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b44c3712-e4fb-4d25-b042-b058b40d0c96 · outbound

This paper cites Approximating.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Approximating

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.692427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.692427Z digest=sha256:0bd68d5a8ca21b3c0e90465279d22076176d3ddf2e1d6d40e21deb99ad0dbba1

Observation 9d978763-eb09-4404-ac3a-ea6d751c60d3 · outbound

This paper cites Thinking Machines Lab: Connectionism , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Thinking Machines Lab: Connectionism , year=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.697244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.697244Z digest=sha256:81e42d9f7b4a7202b993dbdde30d37214f7973482c2f830ff1e4e019b7f7cf58

Observation 3bc3d340-d0ad-46f6-886e-05ad422f23b4 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.701026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.701026Z digest=sha256:ceb5e67223cf7269ad44763d4426a683fe1344220df02fcb4ce80faf46872796

Observation 3f930e07-0f31-4a88-86bd-77b11984480b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.705018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.705018Z digest=sha256:228763aba83fb7215cdf82888086e64ec9b6204adf53f9bdfd7606dc5bccc5d9

Observation 60b9bebc-e6c2-4376-a935-351d30f38d02 · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.708487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.708487Z digest=sha256:0f6e0466e17b12a1c7a0f827290b2307e533f7b68145b3383731965d6a81d782

Observation b4b0f2f6-ed7b-432d-b120-1e6766403ba6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.711697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.711697Z digest=sha256:56048ae965b778331fea258b2928426507cee9484e571619e3e1db0987a82d87

Observation be70df40-b56d-4ea7-a077-17a1b2023331 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.715648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.715648Z digest=sha256:03e45da044461f2224e3ca21eb41d90d0e9d3880847a5f0ff0bf5e46d9c11956

Observation f0787fa9-408b-4eb8-9531-1affac13a384 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.719267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.719267Z digest=sha256:28108cb6350840214271433f25bccd1a930213aa29465c06d4851c968a9a4b0f

Observation a31bca12-0bbb-4470-bd94-c9265e587d90 · outbound

This paper cites NIPS Deep Learning and Representation Learning Workshop , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation NIPS Deep Learning and Representation Learning Workshop , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.722363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.722363Z digest=sha256:3988732966651299b5bfb24c58b81cb3cae858766bc0cc65e4900969b067598a

Observation 1c332071-41cd-4d27-a618-1b40ac4a1c61 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.725125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.725125Z digest=sha256:706e75182994ff1dbc21cc36dc1263d35e4b749dc2a3148660199bf41a280f14

Observation 66007b7a-c580-4046-af05-70c4019a05ec · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.728125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.728125Z digest=sha256:f9634a84bc675ff646316e524d2b8fc9dce0e22cdfa011a9c1d55c5f4d6cb7a1

Observation 5e772e42-ca45-4703-972f-f65ccf585771 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.731053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.731053Z digest=sha256:b7eb94036edefe9a43f24e84754ab2e82862cb1a502be5f8ceb4d004216ea7e8

Observation 58fb0731-3699-499a-973e-7019f2ddc652 · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.734494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.734494Z digest=sha256:ef768030f18fc46a6ceef9cf57a277aea81a2b6efc0e91100f0c68ece64560b5

Observation 301e8b44-978d-47bb-8ef6-e89e0d9de9b6 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.737879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.737879Z digest=sha256:efe8bdbbb891b0eda257e996662f04db648bc8e2f2829aa4c60532ad7929db85

Observation ea2d3af1-68aa-4862-81bf-184310a1320e · outbound

This paper cites 2026 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.740548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.740548Z digest=sha256:48a172293f7e50e9af3b54aeaa78f805a8bdc218903f41586066d7d4b886246a

Observation 7bffd4d0-e002-4b5a-8291-bc548fcdbf54 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.744202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.744202Z digest=sha256:45df770b2a18f5cf175404bc0c0e53273b3f9c2cc259dd2dafe7010aa5095ac2

Observation 1e160db9-9907-4243-9747-5cd4c6284f54 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.748401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.748401Z digest=sha256:0bbe3bdcb88450e04c3d6b418c890b18260ef383bf2cb037c71fe91a33847b7f

Observation 9981f875-77c8-4e0f-a65f-c41b1036962d · outbound

This paper cites Language Models are Super.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Language Models are Super

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.751873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.751873Z digest=sha256:55b5cf355c177df0cca3cde506b38114788828d7866011e62637cda58a4a8f90

Observation f8609371-1e60-447a-afec-bf88a774d8f6 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.755666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.755666Z digest=sha256:6cd1c6f75055fc812f20fbb1c17c81eb349d842b60b5667b235a5ad4125cd9ca

Observation d1c3a387-c06f-4e35-98e1-48f32b788dae · outbound

This paper cites Proceedings of the 39th International Conference on Machine Learning (ICML) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 39th International Conference on Machine Learning (ICML) , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.759548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.759548Z digest=sha256:9c4e73cad3316ca6a2d1913c258fa91a59170b684a0ea6ce890abb75644cfaf9

Observation 1bb26a02-3136-4743-be38-7b335f5a47f4 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Constitutional AI: Harmlessness from AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.762569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.762569Z digest=sha256:a7da51ec8fefb4f024a4031d278a2ba829a68db914c5b9d9b108418bca03411c

Observation e55ff134-a8c9-4fef-be53-5822847b22cf · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.765987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.765987Z digest=sha256:90cbd4a10f90f31dae7745e32ed8748dd23b2ba4e05e1161b218c91635061e3c

Observation f469fe5c-7c2f-4362-a77f-f9cb0ec5a8f0 · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.769326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.769326Z digest=sha256:1b526d4679f76a5fa9578c6e63b296ee8d780949273491e6f68d1cc4f524d644

Observation 2a04f551-0c78-48aa-bcd0-2e3b1650c66a · outbound

This paper cites Hashimoto , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Hashimoto , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.772522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.772522Z digest=sha256:e20e1eb3cf9f5f1a59511f62cc462e3dbfac53a4787abee1afca246de4a8cd95

Observation 8a5ba99a-6b97-4694-a945-a9535aee84ac · outbound

This paper cites Arithmetic Control of.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Arithmetic Control of

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.775564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.775564Z digest=sha256:9e4b01cba371a17a8a3ae4d6ac7524d66421083cb515da31da3add0856c60a3d

Observation 9c6e0cdb-3ef0-409f-b06d-7d3837124cd8 · outbound

This paper cites Rewarded soups: towards.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Rewarded soups: towards

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.779097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.779097Z digest=sha256:11d8e280e7ed07468d5db5c10a555f7752deeed4fe20d4a782a0b2db02a92e65

Observation 5911d5ed-af19-4c20-99d2-f54f52ae8c3a · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.782265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.782265Z digest=sha256:0d27593a699faf8c20160942b57f3efcfa4dbae98997269e0343f5693acb2023

Observation 55315f55-352e-45fb-92d4-c8072d293dd1 · outbound

This paper cites Back to Basics: Revisiting.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Back to Basics: Revisiting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.785597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.785597Z digest=sha256:2e5d27adbe6083fb0aa8dffa7ea47be6638308578cc913e73768f8350db90921

Observation e845bf90-36b6-4458-8001-5b3997453d2c · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.788532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.788532Z digest=sha256:260cc830566ce82cd5ec9739cd3392e0a1b5b9ff2ae4b6677b330d55193db585

Observation f21dcfbf-6f20-4b6c-b3b0-a49a4d561e47 · outbound

This paper cites 2606.16771 , archivePrefix=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2606.16771 , archivePrefix=

Reference 30

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:57:54.687922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.791392Z digest=sha256:08eb66c12f3bc820a9e058c32840c633969bd77811fe30cfacfd02d39acf3fa5

Observation a56a4a33-f71d-4464-8afd-e78f8f8071ad · outbound

This paper cites DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:57:54.634427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.794931Z digest=sha256:81ee7dee142b10d80d177c156276833f230b2de1bf1f412a81c1b62f50c7b743

Observation 947a9e8c-6c31-4e97-b98e-94b3b52ea8b0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.799215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.799215Z digest=sha256:f98e266a11b9c6689fb2264b1c61a2e38d06eb4850b714422080dc5bf80038af

Observation 56257440-42c7-4d8e-8768-c59fd0c926f5 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.802430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.802430Z digest=sha256:eedc385ba788b19619036c6e3d23bac6a7d339e5390e96b46f1cc49cc69c3f2b

Observation 60362f9e-c83b-4f2e-b4c4-f50e93c95758 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.806791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.806791Z digest=sha256:62d9a89a9d1298ecc1d257c884b3f8db8049e18662928fd495bfa2aa180fa3f6

Observation 1355564d-ac98-4810-a57f-0b906cb5eace · outbound

This paper cites Gonzalez and Hao Zhang and Ion Stoica , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Hao Zhang and Ion Stoica , booktitle=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.809737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.809737Z digest=sha256:b1f4475656edc46f29708ba182a656ca14d527f685beb1b7d6542c16f6767791

Observation 2dab2f60-a705-448b-a87e-589b931b20b9 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.813058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.813058Z digest=sha256:d717dd3a23edae00e57707fd149f39ef47f8d8053c9201f897d9e620ca1add98

Observation b3971191-ef85-4ff8-b1c6-f11cee07edff · outbound

This paper cites Patil and Tianjun Zhang and Xin Wang and Joseph E.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Patil and Tianjun Zhang and Xin Wang and Joseph E

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.816389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.816389Z digest=sha256:99f3b3b898a206c149a3e22eba6f61c57dd49d3bbad3c91e03122894fa0444c9

Observation eb7f4ffe-1f1a-4a0c-9444-fbd9de0231a0 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.820297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.820297Z digest=sha256:55a54d09667af344508b2b338a1991b935300b6037cbd33a6d05f71e853c8495

Observation d237aab3-d2f0-45e2-8d98-a83959f16682 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.823262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.823262Z digest=sha256:dccde9cdf681fd82087525cc646c537dfe4f5a49575b6b7e6e4587ce0aab0079

Observation a9c73846-1a10-4f50-9319-e3fb17d0e450 · outbound

This paper cites SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:57:54.619457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.826235Z digest=sha256:9a4d4de85f57056315ea10af499789f61abbb785fab807772b338db26f56769b

Observation 5f1707d8-0451-4278-be2d-ea5f5f65f148 · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.831378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.831378Z digest=sha256:9e05c401cab9f13c20d67eb6c83aee964de45a867cba644f4542162cbf8b8596

Observation 7100e852-0073-42bc-9aef-0a21d22ff9ef · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.836395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.836395Z digest=sha256:ca4d3ddbddc79e7a00d71a8a79c4dac9367534bdd5d6b44a05ea287d5b9b4d85

Observation 458571d0-4753-49b3-bc5e-b90ebef85fe3 · outbound

This paper cites WARP: On the Benefits of Weight Averaged Rewarded Policies.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.839991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.839991Z digest=sha256:694ae6cf2f65f0a668293118b5612096b5e16f3b8982d8fc973d59d6854dec88

Observation b74d5406-0524-46d2-8a75-34d0a3b1ce8c · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: ACL 2024 , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.843008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.843008Z digest=sha256:56f5f4de2fa18d676850edaf07e83c0c99e8221c4323ea5b4ac082ae94c54e8d

Observation 52d38c34-105f-4271-84ec-e8728d765d87 · outbound

This paper cites Singh and DJ Strouse and Tuomas Sandholm and Ruslan Salakhutdinov and Anca D.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Singh and DJ Strouse and Tuomas Sandholm and Ruslan Salakhutdinov and Anca D

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.846073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.846073Z digest=sha256:8ebb57f698021a4f445e5d594e230ce47852128f1367d0a277c180f1cc4290cd

Observation 4a920b9b-be2f-48cc-8f55-8369bda27d7d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.848979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.848979Z digest=sha256:4d72db1ce62dd6b6a82a4cfd30c5f857c0d7e752e42c1367283395cf1ae8945d

Observation 42b17cf8-95dd-4af2-a4f0-cf5d12562fa7 · outbound

This paper cites 2025 , doi=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , doi=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.851928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.851928Z digest=sha256:e4f860473b46d598cecf11519a6ddb2a040ca661e2d84fcdc4701fbfb62cb3b3

Observation 910b124d-5a82-4792-b21f-92dba7479cf3 · outbound

This paper cites 2021 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.855105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.855105Z digest=sha256:a7b4c63b1d08a4fdd5d8aafcef208440fe0260e9872b723a03bc2adcbf3f24ce

Observation b66bf340-d024-4cef-9b55-5130159e6940 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Measuring Mathematical Problem Solving With the

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.858442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.858442Z digest=sha256:92c43dd289b4f4c1838eefb90a24aee4994277fae1edc3ee0bd80a6634376bdf

Observation 7d649c44-756f-4c84-8759-1a43d032246e · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.861265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.861265Z digest=sha256:ec51f3eb7e8eb5a0f5fcaa886a19cc5f31fb26761113efbaacd07da5f1b06cb2

Observation 739164c7-0b3f-434f-b1dd-9d7647fb01c4 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.864101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.864101Z digest=sha256:05dd33615bafc20b31b11ef7807e999a4bd14f3defbee44067d5ee8f07bd18bc

Observation 933ed797-10b5-4678-bba9-08d2b4935b00 · outbound

This paper cites International Conference on Machine Learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Machine Learning , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.867436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.867436Z digest=sha256:040ca9e94debe763b00a38ebd2c9717bb340a0b2cfc356fa1c210da3b1118898

Observation 4c6525fa-6a6e-43e7-b798-c2b3bd17038b · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.871011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.871011Z digest=sha256:7b5cd56086235a9806d83c1d2acf8d7e35512c31d436455f9a1728b794040086

Observation 53692ec1-9980-46a6-ac36-d69d18c923d9 · outbound

This paper cites 2026 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.875132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.875132Z digest=sha256:1f81f752316dadabf384562443e822097f22bd2b9e3252eae78d87b0ccae074e

Observation bc3c0ad2-cb95-434c-b0b0-3abcff808756 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.878335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.878335Z digest=sha256:a87e59e88e49a49bc67f7ca6beb30586f0ca18efe38ca1c3adaf4a60a533f311

Observation 1506ac23-0a71-46b5-a02c-d10d8f928cfd · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.882201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.882201Z digest=sha256:93f8f46ffc19a715b1b85ba9c5b1178bf524347f45d65ad55e811f43def4f9d4

Observation d606a231-f2ae-4509-8597-e32f247ea1ee · outbound

This paper cites Gonzalez and Ion Stoica , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Ion Stoica , booktitle=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.885627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.885627Z digest=sha256:e81fc41c20a87f7a76f084a533ac0021aa8ceda5e3899b44f936cc372f8aaec3

Observation 2370deaa-773b-479e-b087-935cb248bd00 · outbound

This paper cites Machine Learning , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Machine Learning , volume=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.889122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.889122Z digest=sha256:9447b83ad078beccf2cec222515983ba7353a042921edea738f7324093219f66

Observation 4003c66c-0810-457c-b427-a3be5aa1dc7d · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.892306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.892306Z digest=sha256:e1ea91764fb1390ffecb432ab42607b5f347c181534be5e13888ed1be7476341

Observation 96673f63-9191-467a-80d1-a86c82d45f75 · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2024 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the association for computational linguistics: EMNLP 2024 , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.895395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.895395Z digest=sha256:7103563e4f77d3e311d94782e2f80bcc32de4374b9d7dd87413210fe921ba406

Observation fb436b5f-bffc-44f3-8219-b28e115e373f · outbound

This paper cites Everyone Deserves A Reward: Learning Customized Human Preferences.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Everyone Deserves A Reward: Learning Customized Human Preferences

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.898499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.898499Z digest=sha256:08bd2987ea0c5f93cf0f5af42f565b860dc415d3e0d36e838e36d113484f3a27

Observation 17f84292-f9a6-42c0-938a-43315dd4d88a · outbound

This paper cites Kimi-VL Technical Report.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Kimi-VL Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.902060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.902060Z digest=sha256:d76eb1b6da55ec528ec0beb0678ee6dc110586ff9df0e3d9d5368e6b57f3105e

Observation 5c547009-05c7-4ae9-8842-3864d34f42c6 · outbound

This paper cites Advances in neural information processing systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.905649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.905649Z digest=sha256:74fafb5e3d1dda3c433046622561834254b9e5af85290e686e5dc49185d292cb

Observation e7d3048e-1d8c-4e85-be5b-7b2308f1e963 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.909368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.909368Z digest=sha256:192903f355ee9326072bf880b590228ea7087b6552e01f8ce2dfa042746c3d2d

Observation 37c9e7be-5e37-4de4-a6ae-018e2e77da91 · outbound

This paper cites 2000 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2000 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.912159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.912159Z digest=sha256:1f84e82005f5195a6db33f8811f24d7c940baead6a98697a9c7f41a62e317666

Observation ed62110f-f4e8-4f0a-8f43-9a5398d88936 · outbound

This paper cites 1997 , publisher=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 1997 , publisher=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.916002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.916002Z digest=sha256:318e6d41e09dd2912d10d550644437c155a8c832e07dba6d77fbf6866e917a82

Observation 0a830bc8-5b09-4916-afbd-fc7199199291 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.919633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.919633Z digest=sha256:4ad0cf089cd325663d12865f40ba4e9acc9712710d398b419f548e09022be764

Observation 74fe22e8-2260-483f-be44-6ee4dc7feb15 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.922992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.922992Z digest=sha256:1f190f631ca9afaa02b63a1ce4112fe5b49a2065fea1a398d6f611e88258048b

Observation ff4ac2d3-d624-4216-bc39-52052533cbe1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Representation Learning with Contrastive Predictive Coding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.926780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.926780Z digest=sha256:ecdcaf71649d33849f5e4234229c5585fe706aa137a7453d27ef79ac6845a0f9

Observation e8595768-a948-4f78-a4aa-69e7e79a006a · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.930464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.930464Z digest=sha256:e02f306c1969f5bcfd8f2adf9b7e78040f1094892a9211494f3e046775811fe4

Observation 3b0a12de-5c06-49c9-b42c-d1066999ad54 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.933734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.933734Z digest=sha256:ed701232e46e5559c0b95116f4e5856a775ebcb10495f9240bddb0905a301429

Observation efceb99c-e966-482d-88be-1bee64ba9100 · outbound

This paper cites Advances in neural information processing systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.937179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.937179Z digest=sha256:a10f78e34a586b1327804f4f7088cedd97300255bb6e43161af8f2adef7cbb85

Observation e1911532-51fd-466a-9c76-2074d31e8c6b · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.940237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.940237Z digest=sha256:14323d85e98cc8c186812681d06b33dbac22fac4912797c3a8c7bb0ebcc44512

Observation 09f95b76-65fa-442f-af2a-decc09c41a21 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Towards Understanding Sycophancy in Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.943369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.943369Z digest=sha256:f7b472f6fc5a8eda1c3443f59a4447bab1d72c00347e0777c66698ff4bef8b2d

Observation f904c68f-6879-4859-8ce7-4640b3670e3c · outbound

This paper cites Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.946749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.946749Z digest=sha256:fa62ba9cf0bf917a96ce2428e5c4fb8dbf27769adfcf7fb8a99953c7b4249f9c

Observation 5ef0411b-074d-420d-b9bb-4b3e3314c8c2 · outbound

This paper cites Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.950455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.950455Z digest=sha256:b2e3adc7c3a494f5dbe6979cf81eb81f5eb25713728b29126d3bb085cb6158f8

Observation 90036455-6a28-4aa9-8e78-d0208b1910d2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.953858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.953858Z digest=sha256:1b5d75fbc7b7fc10fbc11bff586acce5ab7f791e6c278291d3b7323d88b77137

Observation dcdda4da-491a-432d-a3a1-4b8f546c4c3d · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.957789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.957789Z digest=sha256:f08bfd8a53446a9d189e7d4c0edabc685a024894ded066e7ac323aa3812c95ac

Observation d8724f47-7509-4199-9099-2e2d18c147e5 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.961334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.961334Z digest=sha256:1912d3a53063339a0e06d240e55082a893c961ea50d3a1ecb7545d88ea900e51

Observation 783e67f2-7ba2-4956-96bd-cd9d5f43d6f6 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.964603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.964603Z digest=sha256:9f66b735c65c82b5433e5788bb9f6cb86769a6edb1e44e70d7de368bf5547a39

Observation 283518ca-0f74-43ac-9f73-68e4f8110187 · outbound

This paper cites GPT-4o System Card.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GPT-4o System Card

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.967415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.967415Z digest=sha256:5a2186e42b1a39bfdb378f8116417a8748ee800688c149d19953256486a6c00c

Observation c6c8126d-f837-483d-947e-fbd78482a8dd · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.970403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.970403Z digest=sha256:4105f9528b9cc2ad892931948af24ce109508c6fe524ca6f639b5fe77161d769

Observation 81a4ea9b-7e5b-4972-8c5a-333a96a32c0f · outbound

This paper cites 2023 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2023 , eprint=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.973324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.973324Z digest=sha256:4dff8163da145a7726e7ec4023574905c3b1ab1ef0c8be27f4d7aac98e474697

Observation 8f68f78d-087d-498d-8e1d-b1b9609f78ce · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.976058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.976058Z digest=sha256:c49dfedb6ceb970fe5bfa476d66e2d134c9e899802fa5d863971b1f54f36d8c8

Observation cd26a059-5d3c-4d27-828a-4c16ff5ba1a5 · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.979341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.979341Z digest=sha256:a214da282c6240a94bd9a71e512d6d15aa4e9517cbc45fb49dd32f48a3e9d95a

Observation 2c52bc28-2c74-4fa0-9f45-d1f6cf46fb42 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.982198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.982198Z digest=sha256:f51c01992fb7e2af5660fe18b8c306d465317a0b170c3761c2eafdf6054b52fa

Observation 8d7d7669-d088-4acb-b031-e45a426084a8 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.985088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.985088Z digest=sha256:cecb0410523540011f478ddd2e4644f55a77fef910d30c125eb27870b52d2720

Observation 6d1257dc-eee2-430c-9915-3f920322fb5b · outbound

This paper cites 2021 , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , booktitle=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.988450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.988450Z digest=sha256:716c9cfb8cfc3e31569ee3f1ba6840c000e71adc7b983b97b735f8c974851a69

Observation 98600c1f-5a41-4af8-ae8b-0c0c8d4267af · outbound

This paper cites 2019 , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , booktitle=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.991165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.991165Z digest=sha256:277b6a7068796b70cca1c251dbde792866693c03bfb7f2aab1f74c5643898600

Observation eaeff446-e299-471c-bef7-f32774ea6311 · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.994950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.994950Z digest=sha256:09baa83ac88c92963309769ecb880cb06d04c928e2aee6e96cfc43de82fcff38

Observation c9157bab-d55e-46d8-972f-18644528a140 · outbound

This paper cites The method of paired comparisons , author=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The method of paired comparisons , author=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.998943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.998943Z digest=sha256:7bb4c87cdba2d19834bd89cb7a157007d788926ebc98ad04b99b121ca116ba19

Observation 3671c0d6-78d4-49cf-9a2f-e02ec65f551a · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.001963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.001963Z digest=sha256:43a44f4bfaf4a400cee16bc3f20d4519f18165f6f6daa6607dfec5a7e67b2304

Observation 6db5d0e8-e58a-437c-85d2-fc3a6e5e6409 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.005398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.005398Z digest=sha256:90d94bde4e27eff20f48d0339a171097cb4f5e8b62d018440172f244a2c055db

Observation 8a3c386e-3536-4b72-9788-809f86b0f89b · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.008261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.008261Z digest=sha256:465d1b94706adaedb81d878b91513f15bb8f45956c01d56cfd6af541b09ee035

Observation 6661f9bc-b9bc-4d85-a316-01b58d24a03f · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.011160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.011160Z digest=sha256:7d558cea8f78cc9269d8c64679364a91936ebce1bfc2ff1146ab8e7b662dc431

Observation f7c4114e-0800-4c37-9b5e-fc611cca1038 · outbound

This paper cites Instruction Tuning with GPT-4.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Instruction Tuning with GPT-4

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.014237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.014237Z digest=sha256:4e59eb944c03463e99d02aa616403fe6237702a79307687422f6fe6e322b3474

Observation be4702e3-bdb3-4135-b817-af21f4a2d47a · outbound

This paper cites 2021 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.017797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.017797Z digest=sha256:314dd0c1dd5ab3a14ea8fc506958b208704906f7008f69b8cc5126bc97b0bf4a

Observation 1884369b-ba39-47b9-86d7-cc2fd09e02fb · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.021547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.021547Z digest=sha256:783ecf514b21f4f386723d3f8f0e8b7cdb2a09df149509f527d2a6ee6d7736f0

Observation a58c936d-b789-4294-8d0b-8761a1b6fc64 · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.024618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.024618Z digest=sha256:f5c36dfe620231de31502ecae50611c2e7335f433733c8adeda87c3dbe3a4f91

Observation 9d8a1c50-2af2-488e-b1cd-26f9a98fee5a · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.027500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.027500Z digest=sha256:8e31dc07ce95be0577bc496e99dc6a7b763ae05635192692e61bc9c80300ec2d

Pith citing papers

No inbound Pith citation observations are available.