Pith. sign in

Paper Citation Record · LEDGER

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2502.05434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05434 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:28:41.615055Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:05.568274Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:01:08.101786Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 301c011c-aaae-4ba8-a33f-c7d68292fcee · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.418904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.321150Z digest=sha256:e36fe860223a784859db88179ed0bebee01ed49ab36021004ed4b24d110e8250

Observation 16cf7854-2d9d-4833-9a00-279d84c2eae4 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.409655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.325333Z digest=sha256:7113bcac8633f35a3c542815cb49954046ffc8b64c08aff3fa796c40ae44afe0

Observation c5ece5db-d921-4e3e-8e88-15709e1805c9 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.263004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.329421Z digest=sha256:8a013f230e70e666e0245984618af648d9cb2f0d0188b63cd443cf66d0512f92

Observation 7f48ac4f-09c5-4e78-8c4a-96a53d977655 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.065511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.332925Z digest=sha256:ea136f12da186b48ff884df1323439cef9fc08f504018176ce9e0e5d781f8fba

Observation 5006d36b-43e7-411c-99bf-a58c8d7a1fb6 · outbound

This paper cites For the first term in Eqn.(A.3), using the basic fact thatA − λB/2 ≤ A2/2λB for B, λ≥ 0, we have Et V eE ∗ t 1,π∗ E (st.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling For the first term in Eqn.(A.3), using the basic fact thatA − λB/2 ≤ A2/2λB for B, λ≥ 0, we have Et V eE ∗ t 1,π∗ E (st

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:42.001984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.336619Z digest=sha256:d72df6e16cfdd512e9b8cf83e68b81b38e6b5a41543d45eeb0468beed657ffc9

Observation 52c36661-91b5-4215-8ada-ab2a6923595c · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:41.991949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.340457Z digest=sha256:c498d9343563b25142128b4081bfeb7a57fb834cae34f7bae8ef7dd6cb5d44ae

Observation 022b7e68-6ada-4e95-bfc5-46b5307b7e49 · outbound

This paper cites (A.4) 21 where we introduce the tool ofinformation ratio Γ πt TS t for ease of analysis.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling (A.4) 21 where we introduce the tool ofinformation ratio Γ πt TS t for ease of analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.981743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.345195Z digest=sha256:57e2ae966907e0fe71090a385702e4762d8f2c39dcdfb73fc7a2798a63f6fc11

Observation e3e53918-6140-4950-99db-7c2cc4ed6cdd · outbound

This paper cites However, Lemma C.1 can only be applied to handle the difference between two value functions with the same policy and different environments, while inV eE ∗ t 1,π∗ E (st.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling However, Lemma C.1 can only be applied to handle the difference between two value functions with the same policy and different environments, while inV eE ∗ t 1,π∗ E (st

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.969658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.489650Z digest=sha256:34735524befeecb0cdab5ee97fd6753e3edade1e87b1e627c4939d7584d09782

Observation 292e9b55-a7f7-4a87-8aa1-1acea5ebc1c3 · outbound

This paper cites unifying.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling unifying

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.893703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.532321Z digest=sha256:35f9e02886459d5aba03c5c251f03bb505e4530e4d7f40cd20ff303298e7d729

Observation bd7d0244-7710-47ef-bcb9-7162a8cc4884 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:41.686031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.592222Z digest=sha256:9ef81038a30df73c899ccc321a3093a08368839d786843a35c82a0b89e9629dd

Observation fa6fa785-4cdc-49de-b1cc-a9a01032805b · outbound

This paper cites integrated.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling integrated

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.654315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.610662Z digest=sha256:6c2339c31d879500bb45c714c4d4547b7ee68422b9b95589879a86c8e2958621

Observation 650aaa96-bb07-43e6-84a4-86b62ce39d17 · outbound

This paper cites Since K(ϵ) ≤ ( 3H 2 ϵ )SAH · ( 6H 2√ S ϵ )H · ( 3H ϵ )SAH , 27 we have log(K(ϵ)) ≤ SAH log 3H 2 ϵ + H log 6H 2√ S ϵ + SAH log 3H ϵ ≤ 3SAH log 6H 2√ S ϵ.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Since K(ϵ) ≤ ( 3H 2 ϵ )SAH · ( 6H 2√ S ϵ )H · ( 3H ϵ )SAH , 27 we have log(K(ϵ)) ≤ SAH log 3H 2 ϵ + H log 6H 2√ S ϵ + SAH log 3H ϵ ≤ 3SAH log 6H 2√ S ϵ

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.644467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:28:41.615055Z digest=sha256:5695ac011320df11ee4af9595a99de7551aa2ed71febc22cd80c56fe85c5478a

Pith citing papers

Observation 301e7a58-e6e9-434e-9935-0ccf2266b1bf · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:01:08.191682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:01:05.568274Z digest=sha256:b6139719751d94377cf4cc8bf1fef1f5f85960e473efb6ea04a65fcad51f2a52