Pith. sign in

Paper Citation Record · LEDGER

TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2407.16574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.16574 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:59.576255Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:14:10.944002Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8a80d37e-8a54-43cf-8c36-634b7706e89c · inbound

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning cites this paper.

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:59.576255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:22:59.576255Z digest=sha256:aec86f1f0fef0b50fbe5fe7d22c19a02b4a698b8fe376e9e441d77d7b3106505

Observation 70da07b1-0cf1-44f1-ba5c-0c3613c1b6ab · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.365581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.365581Z digest=sha256:774c5e704fb34c222c23996da69e40c32d0e4c23994c0647a27656824b6419a7

Observation b795ecb2-01f1-4d2c-9ace-c13414b4b50b · inbound

SGPO: Self-Generated Preference Optimization based on Self-Improver cites this paper.

SGPO: Self-Generated Preference Optimization based on Self-Improver TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:49:12.744748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:49:12.744748Z digest=sha256:87b40b07ddac636c870f19e6405b0f8b40d3076b930fe1560bbfca3ea0333bfc

Observation cf5c47bb-8e85-47b3-a615-562ec3d23ee3 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.945445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:e4f86d2114991d383e66a7964411f7ea3a4aded39f6f695bb4b5026af0a1ad44

Observation 9426f27c-1d05-4acf-8f71-d3236dfb0862 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:46.007242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:46.007242Z digest=sha256:495f15be32f3535075b147a413855acc9c29ca62e10c04280e63f62188bb83be

Observation 69356ac8-eeec-4bed-9ebf-82ed49ed2f7a · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:07.503407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:07.503407Z digest=sha256:98026d84ccda64160d0954ac4aa763e3bc67f61831c222c58c71988d99514039