Pith. sign in

Paper Citation Record · LEDGER

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

As of 16 August 2026, this Paper Citation Record lists 100 of 167 outbound references and 0 inbound Pith citation observations for arXiv:2608.03092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03092 v1

Coverage vector

measured 100 of 167 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:54.027500Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 167 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b44c3712-e4fb-4d25-b042-b058b40d0c96 · outbound

This paper cites Approximating.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Approximating

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.692427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.692427Z digest=sha256:9687cf8c28fc93b87234c1108fcb3db8f766727445ea3f895f10d295b8a67b12

Observation 9d978763-eb09-4404-ac3a-ea6d751c60d3 · outbound

This paper cites Thinking Machines Lab: Connectionism , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Thinking Machines Lab: Connectionism , year=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.697244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.697244Z digest=sha256:7291c14ac07a4b9bd6fef62332e193f20e0c570511ec14a02ba3090b50b80b9b

Observation 3bc3d340-d0ad-46f6-886e-05ad422f23b4 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.701026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.701026Z digest=sha256:bfc88172a7e2a81a728f1c93c54930f1f39e4720e68aa99fd36f92fa35b8a543

Observation 3f930e07-0f31-4a88-86bd-77b11984480b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.705018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.705018Z digest=sha256:c15e366aee0f8797f67990a0a902531ba1efeb067c4be964f86378a534f5741d

Observation 60b9bebc-e6c2-4376-a935-351d30f38d02 · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.708487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.708487Z digest=sha256:72255bda4a84b3fe28cfc512e7c3b32cc26b7fb14f0b9f0f2ddfe3769bf87a88

Observation b4b0f2f6-ed7b-432d-b120-1e6766403ba6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.711697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.711697Z digest=sha256:33e7606d77bda0131409c1f0885a09a3ba6f07557fcf98dbe2daba4be8d9770b

Observation be70df40-b56d-4ea7-a077-17a1b2023331 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.715648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.715648Z digest=sha256:6e5b0d113e4878af9d80cc0d26fa91aebd815e1ff3119c9ed7a6055aadb05cd2

Observation f0787fa9-408b-4eb8-9531-1affac13a384 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.719267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.719267Z digest=sha256:741f16dd461e451b54cba609c12692c75b95e7f3bbce5dd7fec69bf20e2c22b4

Observation a31bca12-0bbb-4470-bd94-c9265e587d90 · outbound

This paper cites NIPS Deep Learning and Representation Learning Workshop , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation NIPS Deep Learning and Representation Learning Workshop , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.722363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.722363Z digest=sha256:fdd3eefa9bd2da01d7267a2167d78f543080fe9eedc7a6aaa0a93c1c0f14a417

Observation 1c332071-41cd-4d27-a618-1b40ac4a1c61 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.725125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.725125Z digest=sha256:4a661539596c9184b6d3e3784167d3e56e22c34ad827d6e725d2b3546a279752

Observation 66007b7a-c580-4046-af05-70c4019a05ec · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.728125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.728125Z digest=sha256:39fff611b31e02002682b6f4240a6c0c678b0371e6fb69a8df650ea06ad41124

Observation 5e772e42-ca45-4703-972f-f65ccf585771 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.731053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.731053Z digest=sha256:7e9a84e410641b8d79903a5eb61ee5f8e4c1ccfcfeb51c9d95233be6fa4407b1

Observation 58fb0731-3699-499a-973e-7019f2ddc652 · outbound

This paper cites Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.734494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.734494Z digest=sha256:8c743b19792ffb2c1d656db5ec5b5ce209689b79c476517e563a88607d84cfd0

Observation 301e8b44-978d-47bb-8ef6-e89e0d9de9b6 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.737879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.737879Z digest=sha256:dc49b9c298b5029c23feea78a9ec986aa0309fec8a25e451270a665ee962066c

Observation ea2d3af1-68aa-4862-81bf-184310a1320e · outbound

This paper cites 2026 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.740548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.740548Z digest=sha256:874c0dc636db66e3a9cc5b407c6f94eadbdfa3657c3f7c1656de4f231a270996

Observation 7bffd4d0-e002-4b5a-8291-bc548fcdbf54 · outbound

This paper cites MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.744202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.744202Z digest=sha256:5c371b33e2a34559d7c75c5731fc8fb286e5a314073d84a93dc26344e43351ef

Observation 1e160db9-9907-4243-9747-5cd4c6284f54 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.748401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.748401Z digest=sha256:57af472b570199fab75b50cfdbaf629955405555fa6c749b9b3aaa221cb7d364

Observation 9981f875-77c8-4e0f-a65f-c41b1036962d · outbound

This paper cites Language Models are Super.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Language Models are Super

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.751873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.751873Z digest=sha256:a30c111dd14f43f82349c7a9101c88b056766994d3ce74c8221b812fe940fc73

Observation f8609371-1e60-447a-afec-bf88a774d8f6 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.755666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.755666Z digest=sha256:028b0c23445dafa38b2f2e94f5736a29b5a42b4934ae886af5041f19daf648d1

Observation d1c3a387-c06f-4e35-98e1-48f32b788dae · outbound

This paper cites Proceedings of the 39th International Conference on Machine Learning (ICML) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 39th International Conference on Machine Learning (ICML) , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.759548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.759548Z digest=sha256:70c584ffda1eeabe9a4754015574db300f63e694fa4da1ef3c5bdf8c32238d45

Observation 1bb26a02-3136-4743-be38-7b335f5a47f4 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Constitutional AI: Harmlessness from AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.762569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.762569Z digest=sha256:2ba40b363a119fb446ae26aa88bc141336b2dff7c1d2337bc8be61fa363adb83

Observation e55ff134-a8c9-4fef-be53-5822847b22cf · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.765987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.765987Z digest=sha256:8d1fea524aed98f593f3053cf780ff6c50a490d4565a101927c39a68fec32117

Observation f469fe5c-7c2f-4362-a77f-f9cb0ec5a8f0 · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.769326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.769326Z digest=sha256:604386b27f513f95bed7815cc12bde6bd463eb09e1d9cf5c42bb1866fe353958

Observation 2a04f551-0c78-48aa-bcd0-2e3b1650c66a · outbound

This paper cites Hashimoto , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Hashimoto , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.772522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.772522Z digest=sha256:1c0453d696502ad394332066d9223cd4653d51dc12a837504f1d16d44fc800c8

Observation 8a5ba99a-6b97-4694-a945-a9535aee84ac · outbound

This paper cites Arithmetic Control of.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Arithmetic Control of

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.775564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.775564Z digest=sha256:2deb482ca705c037d2cf4c54ba53419e80b8508acbfcc67ba192c6579b2abc28

Observation 9c6e0cdb-3ef0-409f-b06d-7d3837124cd8 · outbound

This paper cites Rewarded soups: towards.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Rewarded soups: towards

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.779097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.779097Z digest=sha256:ce92555dda4d70b604c2c671d51c5ee81f361a8e96964b8a51921787ed8c0fea

Observation 5911d5ed-af19-4c20-99d2-f54f52ae8c3a · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.782265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.782265Z digest=sha256:801b32b41066016d8a2b9a563695affe5064414905fa60c76beb74702e1b176c

Observation 55315f55-352e-45fb-92d4-c8072d293dd1 · outbound

This paper cites Back to Basics: Revisiting.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Back to Basics: Revisiting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.785597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.785597Z digest=sha256:856bce5265eccdf2112ca56284838ce9517cf6da6d10a0cbb214fc11ca285339

Observation e845bf90-36b6-4458-8001-5b3997453d2c · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.788532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.788532Z digest=sha256:7839714035c40e6d21bd5b2e8e2d62f6dda6988f6266d3e83490a195e057fe3e

Observation f21dcfbf-6f20-4b6c-b3b0-a49a4d561e47 · outbound

This paper cites 2606.16771 , archivePrefix=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2606.16771 , archivePrefix=

Reference 30

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:57:54.687922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.791392Z digest=sha256:5612167bc35f670fe90fb06416a1afb406335bcf16da8c94c7798e800fc90c66

Observation a56a4a33-f71d-4464-8afd-e78f8f8071ad · outbound

This paper cites DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:57:54.634427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.794931Z digest=sha256:40f2f7e40baed4cb3f9d4106ec4e5b9ca098bf90fd8634437466478e065e8dc7

Observation 947a9e8c-6c31-4e97-b98e-94b3b52ea8b0 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.799215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.799215Z digest=sha256:633e8f60a71990f0f94ab17d95610e6c2f9747c9bb9181af6988ffb4b12430f1

Observation 56257440-42c7-4d8e-8768-c59fd0c926f5 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.802430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.802430Z digest=sha256:ae2c874abca0b7f1a3a247660e7c8128c8343700c7ef724a6c02beca2ab06459

Observation 60362f9e-c83b-4f2e-b4c4-f50e93c95758 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.806791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.806791Z digest=sha256:5d08de9a0a127256dd5334d00ada4a0d395634542d8d7e01e0089c76b84aa80a

Observation 1355564d-ac98-4810-a57f-0b906cb5eace · outbound

This paper cites Gonzalez and Hao Zhang and Ion Stoica , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Hao Zhang and Ion Stoica , booktitle=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.809737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.809737Z digest=sha256:8ddb97854cb69f06226efe5b6036e6e296d835db59b1e3d9ccd6b68714946e7e

Observation 2dab2f60-a705-448b-a87e-589b931b20b9 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.813058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.813058Z digest=sha256:5594cf660eee6dba1ba9bfcd1dde6cfc34639e09570a0d2d5e6bcc6c6cfa5bfe

Observation b3971191-ef85-4ff8-b1c6-f11cee07edff · outbound

This paper cites Patil and Tianjun Zhang and Xin Wang and Joseph E.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Patil and Tianjun Zhang and Xin Wang and Joseph E

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.816389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.816389Z digest=sha256:e63946bc8d1832922666c9c2c1875ef62310c371f1754d5437baa014ad89da83

Observation eb7f4ffe-1f1a-4a0c-9444-fbd9de0231a0 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.820297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.820297Z digest=sha256:0d55b84a34f18b37f4ae2aa0f50792f53dfbcb28b93167d3f7d110c9e3173756

Observation d237aab3-d2f0-45e2-8d98-a83959f16682 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.823262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.823262Z digest=sha256:162f61584b52248003d1552b81ee30fb953e512bb242d3ab11a9c43f10abf668

Observation a9c73846-1a10-4f50-9319-e3fb17d0e450 · outbound

This paper cites SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:57:54.619457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T14:57:53.826235Z digest=sha256:2195dddc2739cbbb119d4bbd5949beab915583b6ac7a4ce7d7b6721686dc430c

Observation 5f1707d8-0451-4278-be2d-ea5f5f65f148 · outbound

This paper cites The Perfect Blend: Redefining RLHF with Mixture of Judges.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The Perfect Blend: Redefining RLHF with Mixture of Judges

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.831378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.831378Z digest=sha256:a660e94e5fa63fcdb161746230d61beabe378c7c73fc45b8859621025d35f36a

Observation 7100e852-0073-42bc-9aef-0a21d22ff9ef · outbound

This paper cites International Conference on Learning Representations (ICLR) , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Learning Representations (ICLR) , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.836395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.836395Z digest=sha256:9e8257172819c41801c216fda447dacf1292a27b3b55aa2d41d6a8b4d2c622d4

Observation 458571d0-4753-49b3-bc5e-b90ebef85fe3 · outbound

This paper cites WARP: On the Benefits of Weight Averaged Rewarded Policies.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation WARP: On the Benefits of Weight Averaged Rewarded Policies

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.839991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.839991Z digest=sha256:5fa95405b3874238ce536cce9ea3daa15ee0c559d198a562609b3ee121d07621

Observation b74d5406-0524-46d2-8a75-34d0a3b1ce8c · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , year=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: ACL 2024 , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.843008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.843008Z digest=sha256:1a0dd6e85e130aba9c2590b068b026dbd970fa34335a0bd31250594e1068052b

Observation 52d38c34-105f-4271-84ec-e8728d765d87 · outbound

This paper cites Singh and DJ Strouse and Tuomas Sandholm and Ruslan Salakhutdinov and Anca D.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Singh and DJ Strouse and Tuomas Sandholm and Ruslan Salakhutdinov and Anca D

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.846073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.846073Z digest=sha256:90ef094c82b21ba25ccecbc35b97eced3058037cabed9a5d7f9a139bb949ab66

Observation 4a920b9b-be2f-48cc-8f55-8369bda27d7d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.848979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.848979Z digest=sha256:752a0b4342868657f59bd6c61ddfba66705f20fe920ae562cb97a2ea06477d4e

Observation 42b17cf8-95dd-4af2-a4f0-cf5d12562fa7 · outbound

This paper cites 2025 , doi=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , doi=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.851928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.851928Z digest=sha256:e081e9500c7509ff06169c09cf7747e1fb9156741cbec2dbe730daf3c3efdaba

Observation 910b124d-5a82-4792-b21f-92dba7479cf3 · outbound

This paper cites 2021 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.855105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.855105Z digest=sha256:f00e363536bcb29f3c28fd91b4ae93e245d303021cd484814601d0ee597c12a8

Observation b66bf340-d024-4cef-9b55-5130159e6940 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Measuring Mathematical Problem Solving With the

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.858442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.858442Z digest=sha256:511f3b945becf4af35009400d336402f920b7cc9f9ee4c578e7d20169d4273a8

Observation 7d649c44-756f-4c84-8759-1a43d032246e · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.861265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.861265Z digest=sha256:8d0be7b31597c77700f83098af4bb80538e929b9234e2b42788274709d73d11c

Observation 739164c7-0b3f-434f-b1dd-9d7647fb01c4 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.864101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.864101Z digest=sha256:6683405d139fc0fd41cc8af94b98d3554aa2ccc0e8cbd5059545c22390244a83

Observation 933ed797-10b5-4678-bba9-08d2b4935b00 · outbound

This paper cites International Conference on Machine Learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International Conference on Machine Learning , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.867436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.867436Z digest=sha256:213de20e7be1417af7ffee33c779721e8d040226354e8f4aab00e4877a972cce

Observation 4c6525fa-6a6e-43e7-b798-c2b3bd17038b · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.871011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.871011Z digest=sha256:29dfe77c479ab1301507584866029bb038c2c0668af307b8fc330751222e3f7d

Observation 53692ec1-9980-46a6-ac36-d69d18c923d9 · outbound

This paper cites 2026 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2026 , eprint=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.875132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.875132Z digest=sha256:20e0fabe6ce3b16f66ff4fc3bd1cd02f1fb3808d576d55ccde520578e275a645

Observation bc3c0ad2-cb95-434c-b0b0-3abcff808756 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.878335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.878335Z digest=sha256:6a2971600dbc46ae7b123d85462c6c854af8f944801940c54f046fb11546b049

Observation 1506ac23-0a71-46b5-a02c-d10d8f928cfd · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.882201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.882201Z digest=sha256:611f28517a02e25f6d2367c801bc3a07cc76235962ea894499e086724574bd85

Observation d606a231-f2ae-4509-8597-e32f247ea1ee · outbound

This paper cites Gonzalez and Ion Stoica , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Gonzalez and Ion Stoica , booktitle=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.885627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.885627Z digest=sha256:c78c1c1ef998ccaefc50ef286eb1b3155a0c568a5d5965edf04386360ab0fbd0

Observation 2370deaa-773b-479e-b087-935cb248bd00 · outbound

This paper cites Machine Learning , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Machine Learning , volume=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.889122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.889122Z digest=sha256:6f499b12dd55f53eeb1df7b1d7ebe40b8a25afe8a64a69327fa86e71073453d6

Observation 4003c66c-0810-457c-b427-a3be5aa1dc7d · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.892306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.892306Z digest=sha256:7b7600d56a676e0f197e099f1726d81d6ff91ea9f1558f426e043eac3f2d70bb

Observation 96673f63-9191-467a-80d1-a86c82d45f75 · outbound

This paper cites Findings of the association for computational linguistics: EMNLP 2024 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the association for computational linguistics: EMNLP 2024 , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.895395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.895395Z digest=sha256:1deb9525b8cfdc7655ab45dab46ba42928c5f27613ca2273818d33fd0b1e5253

Observation fb436b5f-bffc-44f3-8219-b28e115e373f · outbound

This paper cites Everyone Deserves A Reward: Learning Customized Human Preferences.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Everyone Deserves A Reward: Learning Customized Human Preferences

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.898499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.898499Z digest=sha256:0a5acdaf457191a7df886c41c3f3e611169676df6ddc5dc0d4e32b19d9439e09

Observation 17f84292-f9a6-42c0-938a-43315dd4d88a · outbound

This paper cites Kimi-VL Technical Report.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Kimi-VL Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.902060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.902060Z digest=sha256:e737f2316341e1382d96cd52b0960d01e3b63ad45b5eb66d326a89745c3de5df

Observation 5c547009-05c7-4ae9-8842-3864d34f42c6 · outbound

This paper cites Advances in neural information processing systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.905649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.905649Z digest=sha256:10cf4842d2a4b98500aaf2cc80082f1f7487bca8fed933e908ba76c0af9a1cb6

Observation e7d3048e-1d8c-4e85-be5b-7b2308f1e963 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.909368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.909368Z digest=sha256:96000317c7abc51b0ab961cdc22aabc7cd86391922024ffbd127b23c1e95ade6

Observation 37c9e7be-5e37-4de4-a6ae-018e2e77da91 · outbound

This paper cites 2000 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2000 , eprint=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.912159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.912159Z digest=sha256:bb7a927ad5cb505fac91a26384c0ca8469550f32c8bc9785646e581cb65924d6

Observation ed62110f-f4e8-4f0a-8f43-9a5398d88936 · outbound

This paper cites 1997 , publisher=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 1997 , publisher=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.916002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.916002Z digest=sha256:3c930cc8da09ad882131e91d978c8d65c112fc5623a6eb0e40af2b457de7cd63

Observation 0a830bc8-5b09-4916-afbd-fc7199199291 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.919633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.919633Z digest=sha256:2830c256ef8abfa06c52795414ad776eb6a3e7e0aa0b70dd16aaf4aedb28e16c

Observation 74fe22e8-2260-483f-be44-6ee4dc7feb15 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.922992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.922992Z digest=sha256:7ab46543764d08e3e78bda18fbc7bf2dcfbb69727671f6b50106788e13f3c2d3

Observation ff4ac2d3-d624-4216-bc39-52052533cbe1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Representation Learning with Contrastive Predictive Coding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.926780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.926780Z digest=sha256:bf34b10a95e9a788ee6816b418b4920c27ba7303a7c9bc0cbc9e965665eee570

Observation e8595768-a948-4f78-a4aa-69e7e79a006a · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.930464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.930464Z digest=sha256:4246d8aa8bc98672d40dbecc781f2971588c38200b2472e85241a676cef0e21b

Observation 3b0a12de-5c06-49c9-b42c-d1066999ad54 · outbound

This paper cites International conference on machine learning , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation International conference on machine learning , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.933734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.933734Z digest=sha256:8b6630b5f168477019907a18644789a7c266f2e54db725a25d62a7be367efaf8

Observation efceb99c-e966-482d-88be-1bee64ba9100 · outbound

This paper cites Advances in neural information processing systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in neural information processing systems , volume=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.937179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.937179Z digest=sha256:87a9be59a528f0acc1a34769dfbdc13f0466edd3bf2b956742d37bf3bc566372

Observation e1911532-51fd-466a-9c76-2074d31e8c6b · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation A Long Way to Go: Investigating Length Correlations in RLHF

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.940237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.940237Z digest=sha256:e6b5b91c4fa50a0ff3fb86c62afba9057023bddcaadfe117ae7fc0434b1148af

Observation 09f95b76-65fa-442f-af2a-decc09c41a21 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Towards Understanding Sycophancy in Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.943369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.943369Z digest=sha256:f643d79f54c24ce0c6992edb84ecd32bde87f8e9987fe06dc48f1de0998ec470

Observation f904c68f-6879-4859-8ce7-4640b3670e3c · outbound

This paper cites Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.946749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.946749Z digest=sha256:c8b13544f63f1154aa34d4f938e704f7966a79121a735b23e49266ff12e0e1fa

Observation 5ef0411b-074d-420d-b9bb-4b3e3314c8c2 · outbound

This paper cites Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.950455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.950455Z digest=sha256:ba49b50690305eaea24d9da29b1d1a437d755e3d64e0c6ddcd0f020a9b5d2a5a

Observation 90036455-6a28-4aa9-8e78-d0208b1910d2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Advances in Neural Information Processing Systems , volume=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.953858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.953858Z digest=sha256:19fea9aedcee4a92197f083b2320953509aa2c787f2d41026647d425460ec813

Observation dcdda4da-491a-432d-a3a1-4b8f546c4c3d · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.957789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.957789Z digest=sha256:4837f574fc6fa3b327abfcb9aa4587118e2481d4a6521f746a3115d76a06db8d

Observation d8724f47-7509-4199-9099-2e2d18c147e5 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.961334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.961334Z digest=sha256:3eea9f82785afa1c7d9a5592e90117caacbf994eae608782301bc27d2600d137

Observation 783e67f2-7ba2-4956-96bd-cd9d5f43d6f6 · outbound

This paper cites an unresolved cited work.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.964603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.964603Z digest=sha256:6241c8b60fb9fa05789416f02e0456d09534fe1055ba49b79650f17367bbf839

Observation 283518ca-0f74-43ac-9f73-68e4f8110187 · outbound

This paper cites GPT-4o System Card.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation GPT-4o System Card

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.967415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.967415Z digest=sha256:7dd6ecb8524adbde4bd3a65099e294b0da92fba8271b919151f3f556252de43c

Observation c6c8126d-f837-483d-947e-fbd78482a8dd · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.970403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.970403Z digest=sha256:f3bf8b1c4043a4cfb7df47cd011972a5f650c8a9779d9acc4c91c645e440b17e

Observation 81a4ea9b-7e5b-4972-8c5a-333a96a32c0f · outbound

This paper cites 2023 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2023 , eprint=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.973324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.973324Z digest=sha256:7683a2db4426184e36e165c36be061938a73846f84dc71bbd81a34d2ca6908f8

Observation 8f68f78d-087d-498d-8e1d-b1b9609f78ce · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.976058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.976058Z digest=sha256:f81a337d1d7a131d52afce5d715d4f77eb786080c2bf32e70f6d5d1ececfd76c

Observation cd26a059-5d3c-4d27-828a-4c16ff5ba1a5 · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.979341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.979341Z digest=sha256:c7c15164c8295217b69714db78410f3f12d91ea0851eb7995209eace32cd190c

Observation 2c52bc28-2c74-4fa0-9f45-d1f6cf46fb42 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.982198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.982198Z digest=sha256:f55afcf56371c5975b32148247d3dee15f1d6d4c69c3540eb55e7edf1d6f84a4

Observation 8d7d7669-d088-4acb-b031-e45a426084a8 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.985088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.985088Z digest=sha256:f8ded02c63a479e06cd77629df0296dc479ccf6f6b7cb401492cebab4b455261

Observation 6d1257dc-eee2-430c-9915-3f920322fb5b · outbound

This paper cites 2021 , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , booktitle=

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.988450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.988450Z digest=sha256:af53ba246065ba59ee1c4d90e8ed7744e51c6c0da8613af8bc51b9ac28609e3e

Observation 98600c1f-5a41-4af8-ae8b-0c0c8d4267af · outbound

This paper cites 2019 , booktitle=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , booktitle=

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.991165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.991165Z digest=sha256:40d9cd92de71629e6da50f8c7cdf71408bd795e0b95d08722b5861d381da7bd3

Observation eaeff446-e299-471c-bef7-f32774ea6311 · outbound

This paper cites 2025 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2025 , eprint=

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.994950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.994950Z digest=sha256:1364f43af4e2304e9cbb278e96a477e33923773f1615ae5a419b2842ec81d2be

Observation c9157bab-d55e-46d8-972f-18644528a140 · outbound

This paper cites The method of paired comparisons , author=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation The method of paired comparisons , author=

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.998943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.998943Z digest=sha256:a1ed85e54ec6a429db1be13eaee435daaad78adfd5fd3f5fb1f0e316960d8e8a

Observation 3671c0d6-78d4-49cf-9a2f-e02ec65f551a · outbound

This paper cites 2022 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2022 , eprint=

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.001963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.001963Z digest=sha256:509207be1be698c5a8fd6d9867139f31b7255101b9058fc4db59aac2a36fca25

Observation 6db5d0e8-e58a-437c-85d2-fc3a6e5e6409 · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.005398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.005398Z digest=sha256:32c4ce0ef24b2c2ce6f0baf42e016b320f7dc91230d67077b7abd93d7e080b4b

Observation 8a3c386e-3536-4b72-9788-809f86b0f89b · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.008261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.008261Z digest=sha256:1d58734a093df136676b7d4a0a71b1debcbf5a670b0ef82d97ec319064b5b3cb

Observation 6661f9bc-b9bc-4d85-a316-01b58d24a03f · outbound

This paper cites 2024 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2024 , eprint=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.011160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.011160Z digest=sha256:e77e6097fd85b4123154c7ef6dcd3e9d91b85da3f09335ef0fabe5888165c7c3

Observation f7c4114e-0800-4c37-9b5e-fc611cca1038 · outbound

This paper cites Instruction Tuning with GPT-4.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Instruction Tuning with GPT-4

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.014237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.014237Z digest=sha256:3c57aba39b9f522da52f3d6d1a97c7367bda5a137fce31bfad9b2d77eb8eb2ac

Observation be4702e3-bdb3-4135-b817-af21f4a2d47a · outbound

This paper cites 2021 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2021 , eprint=

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.017797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.017797Z digest=sha256:fa4a3ac319517911fc3d7cf91f43f4d5fd7a8371b34e6baf157fde84d9722559

Observation 1884369b-ba39-47b9-86d7-cc2fd09e02fb · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.021547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.021547Z digest=sha256:9fba87af06371e9b19fe04977d73661c3ca9a7f37aa0513bf33653f0b3d38c32

Observation a58c936d-b789-4294-8d0b-8761a1b6fc64 · outbound

This paper cites 2017 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2017 , eprint=

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.024618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.024618Z digest=sha256:7f24955bef8da534f667170f2bb12dadeb57ba60bd5ee83fced37bc643993e02

Observation 9d8a1c50-2af2-488e-b1cd-26f9a98fee5a · outbound

This paper cites 2019 , eprint=.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 2019 , eprint=

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:54.027500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:54.027500Z digest=sha256:6c32524409c92f461ab8a7a33a69b6152cb8cc6aee5b9abb08873fd9f2c8b95f

Pith citing papers

No inbound Pith citation observations are available.