Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:45:53.918108Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2505.06601.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:45:53.918108Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0329729c-1293-4c61-85e3-ccfa75fe2e25 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 84cdfbd8-5826-4501-9843-5d8d24bbc0fa · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Now, we define two sets, given anyη∈(0,1), S1 ={s∈S: 0<r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)≤η}; S2 ={s∈S:r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)>η}
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a370af60-ad00-4d31-a9cb-6e3fc2fe9c40 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d2ec498a-836a-462a-8c4b-ca0dec5f7998 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation daf2ee32-c232-43a1-bfa2-15b3c775353e · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b0801597-7ea2-4518-b8f9-03d2e6f84c16 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Without loss of generality, we also define classes of sub-networks ofFDNN, that is,{F1,F 2,···,F |A|}, with non-sharing hidden layers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 11e422ca-cc1a-43b4-9d64-cde127037e15 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 40a15220-f7fa-43f0-886a-384e0e4f9b36 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9e009a17-72c6-4c14-a707-f9738718d352 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks □ Different from the deep regression and classification problem, the convergence of the excess risk does not directly ensure the functional convergence of the estimated reward
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e5792b3a-01cc-4237-af5e-36dad5b3bfd8 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks and Bayati, M
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cf49d02-47d6-4dca-a9cc-20311bf1a2e1 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Dual Active Learning for Reinforcement Learning from Human Feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a897fcd-bf4b-46a4-b36f-645274ee75c9 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks and Kupper, L
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ff05ea61-5cf2-4448-bb54-cb05ad9c6d9f · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dca5193-da27-4c43-a9f9-60b89abe5e22 · outbound
Learning Guarantee of Reward Modeling Using Deep Neural Networks B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M
Reference 1897
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.