Pith. sign in

Paper Citation Record · LEDGER

Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2405.17931.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.17931 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:08:36.108422Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T06:25:27.959448Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9da7accf-d0c6-4a84-83b5-c3497a848cf6 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.597229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:9a291be9c7c492b81f10f6f942e2fa0e9acad96b3c35cb6d30c2d4a04939eb14

Observation f40e3127-5121-4f19-9b5f-901994342f6c · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.962311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:231d03603b0bfc9de729d5daf544640676980ec9239d5fb01662197ab93ce440

Observation 158db7b5-0ecb-4273-b7e3-4f4eb3b32069 · inbound

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging cites this paper.

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:36.108422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:08:36.108422Z digest=sha256:e0a15d5de7bc988b3ab1b54f78e4e6f5729d5c4e9d739fef6cf671062eff06bb

Observation 9d51499f-dddf-4b90-bac9-be034ba83a96 · inbound

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling cites this paper.

Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 240

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:59.798792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T03:05:36.871497Z digest=sha256:ddec1cbf9dfebf6516f225a0f497426e714579da00cc28a9b975a1cb695ea38c

Observation 37544946-5c88-46e8-94cd-9a0dfcf70f97 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 252

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:84f6d8ba06eb1bfa8395309a47dcfc53d31eee88f786b11dcfb2d6c9365d1c8b

Observation 63836574-52a5-499e-bf51-ddf136e7e96f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 253

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:01.653546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:01.653546Z digest=sha256:78567f63fb458f1a3d232c9f4eab095d0c61b1d51651ba577495fae157c2def2

Observation 03c19841-5f0e-4102-9393-f561e2107b15 · inbound

Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning cites this paper.

Relative Parameter Importance in Task-Agnostic Replay-Free Continual Learning Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T02:17:28.447082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:17:28.447082Z digest=sha256:ed7ec898c6b5a5b878f5060074b5f8fb2e87cdc7115fd50334e7cb6f27f52eae