Pith. sign in

Paper Citation Record · LEDGER

Transforming and Combining Rewards for Aligning Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2402.00742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.00742 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:06.839256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:07:47.876920Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3fffca85-55af-4cda-a00a-bed68e84db18 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.399932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:1c09fdb8665aade7e8b1aafcd275e69c114a53adbb7a9da78eb63fbc5ff2e3db

Observation 4909cd43-a8e3-4144-8f5e-775542ef432e · inbound

Language Model Networks: Supervision-Efficient Learning through Dense Communication cites this paper.

Language Model Networks: Supervision-Efficient Learning through Dense Communication Transforming and Combining Rewards for Aligning Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:51:41.910711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T14:50:45.735917Z digest=sha256:03557ba60465d9031e0e2323f750a15dd1a583476a5f1cadce2617dcf007b20a

Observation c7c2a000-22cd-47a4-8c5e-76e767e818af · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.839256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.839256Z digest=sha256:8d884d55badea6a67728cf56c698fc4fd07aaa7fea693a25c26066ef8db01333

Observation 1378a620-b05f-49a2-b875-d8cf129227b0 · inbound

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective cites this paper.

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective Transforming and Combining Rewards for Aligning Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:06:24.003533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:06:24.003533Z digest=sha256:0a03dbec0852b662188f11a80fef418f9c4c5b4ad5c2184a3b454993943f9b6a

Observation cf7406b8-da57-48b7-a5b5-fb213d6d1f82 · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Transforming and Combining Rewards for Aligning Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:12:14.083693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:2e8926cf4d5d28e2c6a5d6125aafca47c27d07e16cd9564790658d4d8fcd51e4

Observation aec05c2e-5af9-45a5-9870-bee120dd3755 · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Transforming and Combining Rewards for Aligning Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.878366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:824def880eae1a70564bcdc0cdb2ee534aa796b7091a77737bb9b88b4e758f40