Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2305.18438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.18438 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:11:32.681632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:10:10.504002Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fab1eae3-f722-4883-acd2-cc5ca4322825 · inbound

A Survey of Research in Large Language Models for Electronic Design Automation cites this paper.

A Survey of Research in Large Language Models for Electronic Design Automation Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:55.570087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:55.570087Z digest=sha256:dd08f79fe388833cd29f88eb40b380ec9e0803e0042458ef40b569ad4a078efe

Observation 80cabdce-42f1-4842-bbbe-f9b87b4f4220 · inbound

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models cites this paper.

Harness Local Rewards for Global Benefits: Effective Text-to-Video Generation Alignment with Patch-level Reward Models Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:26:17.335567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:26:17.335567Z digest=sha256:fa61f58a43ef3b56783ad7f97173e51ca8b38b1c36d5d90cb1c1d674f0c4fb27

Observation e5f631b0-48e9-4dd1-8bde-3ecdd832eefd · inbound

Contextual bandits with entropy-based human feedback cites this paper.

Contextual bandits with entropy-based human feedback Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.751446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.751446Z digest=sha256:497062243d6101881ed9925627e70e2fafde88c6310a1077212bb666d7053c39

Observation 666d76c9-529d-4110-a30d-2ab3e50105bb · inbound

Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing cites this paper.

Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:11:32.681632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:11:32.681632Z digest=sha256:816eec7a51fdc27882a58ac89e7b258f1112e0cefac9a1423b68471d054f3fba

Observation 3da1400d-4fe1-460d-b28c-821007d36f2a · inbound

On the Limitations of Steering in Language Model Alignment cites this paper.

On the Limitations of Steering in Language Model Alignment Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T04:28:15.727690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:28:15.727690Z digest=sha256:81aa04d310c9056d470a49d3fafa3e990ba1de0004438a21779aed470c9acced

Observation 2030cac4-656a-4f31-95f2-c5614a17aaf4 · inbound

PLHF: Prompt Optimization with Few-Shot Human Feedback cites this paper.

PLHF: Prompt Optimization with Few-Shot Human Feedback Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:44.648590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:36:44.648590Z digest=sha256:98ed286b0e5fd28a5a240e677955b90b4ef14c24a675f8654169cfcfb6508d75

Observation 0b1995dd-98b1-412c-a68b-ce9cb862c1ef · inbound

MR. Judge: Multimodal Reasoner as a Judge cites this paper.

MR. Judge: Multimodal Reasoner as a Judge Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:51.799186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:51.799186Z digest=sha256:a54c87618fa4f046d08b041d48462290e43f8b855ab7bb7bd46be9589c63c837

Observation b87d7704-760f-4733-9942-2bbf1c04246f · inbound

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning cites this paper.

Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:27.735776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:37:27.735776Z digest=sha256:8dec883ff8d4517ab198132bd87535f38049985ce52d2258d571b01337d74fd1

Observation 42031251-04a4-4d60-98f6-1d92937d10ec · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.277156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.277156Z digest=sha256:7ed70f5425fc9e5b63439b84479bafd952a49d76ccea4de116baff665476c2c3

Observation c477f12d-8642-45e5-9f7c-6d8ee20d0c30 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.572546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.572546Z digest=sha256:b617dcc7c0e8ad68f605de6d61a085af391aa073be4154f230f3177be45ac480

Observation bdc3fece-0080-4353-858d-1a6b33231a1c · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:16.570144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:16.570144Z digest=sha256:1ecc963b1b07bf93c9120439e4a96e39184fa9a182f0072ec1bcc9fb189c23c7

Observation 8aa7a4a6-07d9-461f-b276-4e1c25f312bc · inbound

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems cites this paper.

MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:28:39.988703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:28:39.988703Z digest=sha256:fa94a7c0ca7fd3a07f0dab4ba717af3dc86f3510a4645ac6ca42c904d05cccac

Observation b3e76895-0dd6-4821-bc4e-f9143da03a51 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.513538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:39b0a0a22ce32bc105b485b11f4cdb981d3cbbcf4ed36205f2956d9a4173ac3d

Observation ae595c80-a756-489f-b7e0-67610fae47ce · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.505952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:e2d2c0fdeb585e30e52a89dcae824f90acd22b3e83c2c0c82117382232609c53

Observation 12e7541c-527f-4475-86c4-d07f24850210 · inbound

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback cites this paper.

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:19:29.068443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T21:03:48.813600Z digest=sha256:50e20febca251571b02b91310918ddc4c89e4d132d13cffb4b675872c1218f4d

Observation d6d6702d-479e-465c-a86c-c1f3457a1d27 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 288

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:99467b3cb7156c0e926885cc8adfdebb87a75a308e8e9eeb1c32e1b38aef1577

Observation df9ce709-e49a-4775-8246-c31f7cc44862 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 289

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:06.266778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:06.266778Z digest=sha256:a9403f8f1810c78adbef244ff7da259bb0286a127238d56d0dabef721d54054c

Observation b95f051c-6a8e-40bc-9447-5dec8e7c6fe0 · inbound

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization cites this paper.

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:12.905158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:12.905158Z digest=sha256:e1233e36c35364b0a9ef2b41a085cd92ae46afa605359b0ae2e3850096539e92

Observation f02a5a78-e000-47de-a26f-568320d40a04 · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.856523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.856523Z digest=sha256:9d79152511ef27f3a055f20ac5d1b0be846da34e4c53625498be5b574390cdf8

Observation ffec75fe-ad00-4c27-ba62-c9b261665e62 · inbound

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling cites this paper.

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 263

Resolution
unresolved
no resolver link, observed 2026-08-12T14:10:46.225203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:10:46.225203Z digest=sha256:ba034cf54deafdbc064180642533e1b170e2539c1f700eb87db88c6123accbf7