Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning from Human Feedback with Active Queries

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.09401.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09401 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:41:30.916115Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:09:19.274569Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7036c7d0-9413-47a1-95c5-49fd2b504709 · inbound

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework cites this paper.

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework Reinforcement Learning from Human Feedback with Active Queries

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:15.598159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:15:15.598159Z digest=sha256:fe3cd9844aafbb20bfedb381056365f4bbd731818ff2bfdfb2161c8f7ca36109

Observation 41d6af29-5e50-4373-a0e9-f826e1ef227a · inbound

Test-Time Alignment via Hypothesis Reweighting cites this paper.

Test-Time Alignment via Hypothesis Reweighting Reinforcement Learning from Human Feedback with Active Queries

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.452328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T06:55:54.051821Z digest=sha256:4c1bc00e8600da0d6df18141d1fbd4e2d48fb72280cb8efd1945f5ff6ad8d05f

Observation ab14e6a3-2ddf-46d5-9cb7-a967ae82c36b · inbound

Federated Linear Dueling Bandits cites this paper.

Federated Linear Dueling Bandits Reinforcement Learning from Human Feedback with Active Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T16:44:30.647246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:44:30.647246Z digest=sha256:8424bcf81d4838e47ac896bfd010adcfe78ac1d1f711e880d96050c39a010ba6

Observation b963b863-fbed-47c6-a40a-e41b21635caf · inbound

PILAF: Optimal Human Preference Sampling for Reward Modeling cites this paper.

PILAF: Optimal Human Preference Sampling for Reward Modeling Reinforcement Learning from Human Feedback with Active Queries

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.095156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.095156Z digest=sha256:063e848534c7488f0c318e47d2a961713f5947033ff42fee6b2c7b78f4b99f11

Observation 19c40104-86f5-48a9-a5f5-d6374f1e5341 · inbound

Active Human Feedback Collection via Neural Contextual Dueling Bandits cites this paper.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Reinforcement Learning from Human Feedback with Active Queries

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.916115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.916115Z digest=sha256:675ef471926078820cc5353a81876191cb1cf6ea8de9b047477ceee46fd28002

Observation c6d2d641-20a7-4abf-a3ef-313a794c812e · inbound

Reinforcement Learning from Human Feedback: A Statistical Perspective cites this paper.

Reinforcement Learning from Human Feedback: A Statistical Perspective Reinforcement Learning from Human Feedback with Active Queries

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:13:13.738953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T20:10:43.578904Z digest=sha256:d99da53ae95d0909c64af3c12d379fca8573debf3f930c55ccc926a6a44b076d

Observation d3d67732-188f-45fd-9f06-97da319aa2c5 · inbound

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards cites this paper.

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Reinforcement Learning from Human Feedback with Active Queries

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.309654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:58:46.313028Z digest=sha256:d995131fea72c2832c8bcc5d08572d262a527d8346383f3ac26179a19576f8e6

Observation 59cdca9d-327a-4ca3-ac54-0d858eaffd4d · inbound

Which Pairs to Compare for LLM Post-Training? cites this paper.

Which Pairs to Compare for LLM Post-Training? Reinforcement Learning from Human Feedback with Active Queries

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.276435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T20:37:19.221848Z digest=sha256:d99e6a9cfea198e33118012fb70e1c6ae4bf75840edd2d570aeeb4ca0728180f

Observation 755b9f24-0268-4a2b-9487-2d017ca150c1 · inbound

Personalizing Large Language Model Agents with Small Policy Models cites this paper.

Personalizing Large Language Model Agents with Small Policy Models Reinforcement Learning from Human Feedback with Active Queries

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:06:56.090015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:06:56.090015Z digest=sha256:caf8a41191c254e800a75a301b57783544252f765b3b16f4ec6a2e019361587c

Observation 0aa0e2de-c844-4422-9fc6-689d858e27e6 · inbound

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design cites this paper.

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design Reinforcement Learning from Human Feedback with Active Queries

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:19:36.565683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:19:36.565683Z digest=sha256:2ab61dabde8dff2c8130a0706827de8002e0828558ad5d57ea6ed35f906ab501