Pith. sign in

Paper Citation Record · LEDGER

Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:1907.12894.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1907.12894 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:25.809592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T01:46:18.556913Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b9edcfba-e7ac-45d7-9864-73a70b203a45 · inbound

Learning to summarize from human feedback cites this paper.

Learning to summarize from human feedback Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:46:18.560471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:46:18.486086Z digest=sha256:c9e6f17b6d735f6a8463681ea4b40aaeabbda0fae90376def5f8c495a5061566

Observation 28d9f7d8-2d84-4c68-8936-eca7a58334a1 · inbound

RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback cites this paper.

RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:32:27.926085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T21:32:27.806494Z digest=sha256:01a6c8c8e22a95b3acc36b8962a3bdc5895e6b07fac41794d746bb8018e36817

Observation 945bfaa9-cd47-4ce1-be80-991b4affb169 · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.554225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:fcc9e50248b5e03263c861b17bf39f9e8bc6218bd379f0727d2e8d09bcf95f9c

Observation dc605c3a-310f-4ff3-9b8f-329b499d21c4 · inbound

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs cites this paper.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.809592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.809592Z digest=sha256:22eca10b17e3deb0088d193cce74681ff751296ed663bf724cf605abe8161207