Pith. sign in

Paper Citation Record · LEDGER

Transforming and Combining Rewards for Aligning Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2402.00742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.00742 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:06.839256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:07:47.876920Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3fffca85-55af-4cda-a00a-bed68e84db18 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.399932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:e039d38fcfdb3496d61b9f1d6f99c2681c50f11221d957bc00e37171b61d2f1f

Observation 4909cd43-a8e3-4144-8f5e-775542ef432e · inbound

Language Model Networks: Supervision-Efficient Learning through Dense Communication cites this paper.

Language Model Networks: Supervision-Efficient Learning through Dense Communication Transforming and Combining Rewards for Aligning Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:51:41.910711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T14:50:45.735917Z digest=sha256:6a843afc635b7a567eb46b9a1714f8940d75a7e97d6c148cd8f6d4c81007b7c0

Observation c7c2a000-22cd-47a4-8c5e-76e767e818af · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.839256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.839256Z digest=sha256:8d884d55badea6a67728cf56c698fc4fd07aaa7fea693a25c26066ef8db01333

Observation 1378a620-b05f-49a2-b875-d8cf129227b0 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Transforming and Combining Rewards for Aligning Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:24.003533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:24.003533Z digest=sha256:0a03dbec0852b662188f11a80fef418f9c4c5b4ad5c2184a3b454993943f9b6a

Observation cf7406b8-da57-48b7-a5b5-fb213d6d1f82 · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Transforming and Combining Rewards for Aligning Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:12:14.083693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:b6d2410ad2eb8989b402125064c363850e315ae809b0feebd1b8d9060b4a5370

Observation aec05c2e-5af9-45a5-9870-bee120dd3755 · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Transforming and Combining Rewards for Aligning Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.878366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:326a8d3abdb9b294fe8339dc7e6c8c6af6fa17a350628ba29e974b19c1de16bb