Pith. sign in

Paper Citation Record · LEDGER

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2501.02790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02790 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:26:34.918099Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.511368Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5923323-4fcb-4b42-9458-093c65a4a7df · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.401522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:52c17c32bee8cb7590d682432ac1cf9d6eee48b9f3970fed80f041a3845f956e

Observation 35917a05-ad3b-432d-881f-e8f9db0b75b6 · inbound

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment cites this paper.

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:18.030277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T23:00:14.031709Z digest=sha256:cd8bbb852054f70912b006e464f81e5589f90c36038f4f1f1124ca324d20398f

Observation 26de573c-73ff-40fa-a7b8-64a60a72c502 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

Reference 151

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.512754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:9f8ac167911fc51a52ee77748ecc783f6f63a0654449358beb643f0c67df1554