Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning from Human Feedback with Active Queries

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.09401.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09401 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:41:30.916115Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:09:19.274569Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7036c7d0-9413-47a1-95c5-49fd2b504709 · inbound

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework cites this paper.

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework Reinforcement Learning from Human Feedback with Active Queries

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:15.598159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:15:15.598159Z digest=sha256:ea3631a461091c87e1ed663e9e2d797813e0729f0cbe399fa91e9776658f4edd

Observation 41d6af29-5e50-4373-a0e9-f826e1ef227a · inbound

Test-Time Alignment via Hypothesis Reweighting cites this paper.

Test-Time Alignment via Hypothesis Reweighting Reinforcement Learning from Human Feedback with Active Queries

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.452328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-23T06:55:54.051821Z digest=sha256:a3bde113597a7f5d209aa4cb778342a0cda8917c136e6fbbcb0f27635f7b3ed4

Observation ab14e6a3-2ddf-46d5-9cb7-a967ae82c36b · inbound

Federated Linear Dueling Bandits cites this paper.

Federated Linear Dueling Bandits Reinforcement Learning from Human Feedback with Active Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T16:44:30.647246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:44:30.647246Z digest=sha256:6ad26f9601aeea0f5c55cb7cbd9c99249f2a234442ea58723e407c2627e62fd7

Observation b963b863-fbed-47c6-a40a-e41b21635caf · inbound

PILAF: Optimal Human Preference Sampling for Reward Modeling cites this paper.

PILAF: Optimal Human Preference Sampling for Reward Modeling Reinforcement Learning from Human Feedback with Active Queries

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.095156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.095156Z digest=sha256:972b61e6981f9bf00b438686384346d72e4d46346ec4896db5c0fc2b6cf4bd2c

Observation 19c40104-86f5-48a9-a5f5-d6374f1e5341 · inbound

Active Human Feedback Collection via Neural Contextual Dueling Bandits cites this paper.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Reinforcement Learning from Human Feedback with Active Queries

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.916115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.916115Z digest=sha256:8af5abbf48ee773fd2a7b85403b4fc38e8f625e906ed7ee8ce0b4b758414a2b5

Observation c6d2d641-20a7-4abf-a3ef-313a794c812e · inbound

Reinforcement Learning from Human Feedback: A Statistical Perspective cites this paper.

Reinforcement Learning from Human Feedback: A Statistical Perspective Reinforcement Learning from Human Feedback with Active Queries

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:13:13.738953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T20:10:43.578904Z digest=sha256:33ed399a67fcfff9b64f9d78b22bf849809edf8ff9bde635e2cb6a0be4e239d4

Observation d3d67732-188f-45fd-9f06-97da319aa2c5 · inbound

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards cites this paper.

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards Reinforcement Learning from Human Feedback with Active Queries

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.309654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T22:58:46.313028Z digest=sha256:174d9d17257a27eb5609f72e6f32b65a2b5acb1750a8226a848a4a91a7af3d20

Observation 59cdca9d-327a-4ca3-ac54-0d858eaffd4d · inbound

Which Pairs to Compare for LLM Post-Training? cites this paper.

Which Pairs to Compare for LLM Post-Training? Reinforcement Learning from Human Feedback with Active Queries

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:09:19.276435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T20:37:19.221848Z digest=sha256:9c7a1f8f1722a7230bf4370b3fb76fce6107f656605eddcef0e4511fe1e30a87

Observation 755b9f24-0268-4a2b-9487-2d017ca150c1 · inbound

Personalizing Large Language Model Agents with Small Policy Models cites this paper.

Personalizing Large Language Model Agents with Small Policy Models Reinforcement Learning from Human Feedback with Active Queries

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T01:06:56.090015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:06:56.090015Z digest=sha256:60c5b13b186d03c0846ddccaf45518865aba8fcda295427dbc3fb1111431741e

Observation 0aa0e2de-c844-4422-9fc6-689d858e27e6 · inbound

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design cites this paper.

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design Reinforcement Learning from Human Feedback with Active Queries

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:19:36.565683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:19:36.565683Z digest=sha256:6b3cbc4a0e13a626d5288ca9f76ae526689a74f4c9520561d2bf803182978072