Pith. sign in

Paper Citation Record · LEDGER

Training Language Models with Language Feedback

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2204.14146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.14146 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:56:20.331588Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T04:06:34.883183Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 63b8fb82-1635-492f-b352-64a83a98501a · inbound

Aligning Text-to-Image Models using Human Feedback cites this paper.

Aligning Text-to-Image Models using Human Feedback Training Language Models with Language Feedback

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:15.534040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:39:15.388881Z digest=sha256:883dbc0053724e2c1bf6a271aa2100d0aafc367d60de3b48ae20a119b02827ed

Observation 122389d0-cc1d-4224-bce7-9441acbf0b45 · inbound

Self-Refine: Iterative Refinement with Self-Feedback cites this paper.

Self-Refine: Iterative Refinement with Self-Feedback Training Language Models with Language Feedback

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:47:39.752051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T20:47:39.476572Z digest=sha256:3d2085ca1f08a1a503c03a2e0e9d6a5c85fffdd8f6472953b920ca45f3cb810e

Observation 3ae649dc-bcb4-4e30-9704-7e0ef0cfd0c6 · inbound

Copilot Arena: A Platform for Code LLM Evaluation in the Wild cites this paper.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Training Language Models with Language Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.331588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.331588Z digest=sha256:fe2d43065dff2a1f27ba8dd796de94f5cdf8c126d6ca47a7bd85be70c7cb33c6

Observation b8e76d64-de31-4412-8cc7-9f07810719ef · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Training Language Models with Language Feedback

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.166368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.166368Z digest=sha256:818ab38bbbcbfd92aafff35b6445231ac0e95237df35e51f8e94428662fde5ae

Observation 4840049b-dc0a-49f4-bd18-1f7e140e7981 · inbound

R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification cites this paper.

R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification Training Language Models with Language Feedback

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:50:29.461048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T07:47:24.092334Z digest=sha256:9a3d26435d7436ff9a7b806d121582a0f3a2cb6359c6c5366d95711090d20584

Observation a56a5119-5298-4e4a-96c2-c2b705a234c5 · inbound

Characterizing initial human-AI proof formalization workflows cites this paper.

Characterizing initial human-AI proof formalization workflows Training Language Models with Language Feedback

Reference 298

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:06:34.884497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T09:29:50.282874Z digest=sha256:58806a49d20b2d2a1f616a135fd2ea6e0eac2d249028148d2b817d8e7db60c91