Pith. sign in

Paper Citation Record · LEDGER

SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2203.10050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.10050 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:00.332041Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:06.249099Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 35e6f199-18ce-4b1f-b74f-06ac3a279846 · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:58:33.443262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:c0739384049fd7a6a603a3d9ce40f345eae7cda08535bea042ccd2c0912cbe20

Observation cba3728a-2163-4aaa-8eee-1e95afce2404 · inbound

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries cites this paper.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.332041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.332041Z digest=sha256:195c28e3a9f5e28fc28051b7165ef615c884bef1e0ba5d21c77e3bab442992dd

Observation c0ed7339-82d8-44be-a7b7-a506a191fd40 · inbound

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning cites this paper.

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:03.800435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:03.800435Z digest=sha256:a5c0354bd103cc1d31e81c7cdb1fbef9086b0166642b048f1761e7113d507f05

Observation 60a89451-5820-4f34-b2dd-015d8b93da35 · inbound

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning cites this paper.

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:31:36.413182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T13:27:48.404574Z digest=sha256:3bb969a7c7ac09891d03cc30b0b67a16318524a8cfab587cb9d6ed55bea0f62b

Observation a33cc633-ab9e-40c0-8d3f-d278959f06ca · inbound

Residual Reward Models for Preference-based Reinforcement Learning cites this paper.

Residual Reward Models for Preference-based Reinforcement Learning SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:36.478826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:36.478826Z digest=sha256:13edf488a1a17be89a4f6358291bd6f02c0ec39dd27edb378e893afd04c903a0

Observation 5a7a11f4-574d-433d-a85e-a8c1ccd9e8f5 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.525502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:4995e7485b7609d09d110f2a8da6bb56eb835d159c75f181ee5a2ca8f2ff2b9a

Observation a642df2a-df81-4add-97a5-b3b176611927 · inbound

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback cites this paper.

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.087977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T23:35:03.577967Z digest=sha256:e7eca5fe1b4cdcede42a9651341c5670cba8d3fbe450df22c5508b9dfbe5f194

Observation 45516930-8444-476e-9e21-69b2c7f7a3fa · inbound

MAPL: Multi-Objective Preference Learning for Robot Locomotion cites this paper.

MAPL: Multi-Objective Preference Learning for Robot Locomotion SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:06.250669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:11:00.753949Z digest=sha256:0e8a41dcfcf9dd1813dd6c8df2d02bf12d5cf6dbf2aab59a01e5af93f0abd6e2