Pith. sign in

Paper Citation Record · LEDGER

Learning Guarantee of Reward Modeling Using Deep Neural Networks

As of 17 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2505.06601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06601 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:45:53.918108Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0329729c-1293-4c61-85e3-ccfa75fe2e25 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.107037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.870651Z digest=sha256:742d2e688c86f25bdff26323e7828e05e61afde686bc4cc11b99ef2bf190bd30

Observation 84cdfbd8-5826-4501-9843-5d8d24bbc0fa · outbound

This paper cites Now, we define two sets, given anyη∈(0,1), S1 ={s∈S: 0<r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)≤η}; S2 ={s∈S:r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)>η}.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Now, we define two sets, given anyη∈(0,1), S1 ={s∈S: 0<r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)≤η}; S2 ={s∈S:r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)>η}

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.046769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.904919Z digest=sha256:6e4f77434031055afc47dff75c881705d0589cdf4085cd6d6c0ad16f1e5463f5

Observation a370af60-ad00-4d31-a9cb-6e3fc2fe9c40 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.097126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.878712Z digest=sha256:c7fcec3a95d98e801764d00fb000fd20480456c21a16fc3411bc950fb73e49d5

Observation d2ec498a-836a-462a-8c4b-ca0dec5f7998 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.067826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.897920Z digest=sha256:22868818cbe2cd9a7185830eb6c1b2c2425983c4f62d9ef5f8e36c01ccc76d29

Observation daf2ee32-c232-43a1-bfa2-15b3c775353e · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.057379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.901469Z digest=sha256:0e04beb777b58ea0697e06013fbb293ce6891dd8c40be3a9ed7ccf13fbf4777b

Observation b0801597-7ea2-4518-b8f9-03d2e6f84c16 · outbound

This paper cites Without loss of generality, we also define classes of sub-networks ofFDNN, that is,{F1,F 2,···,F |A|}, with non-sharing hidden layers.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Without loss of generality, we also define classes of sub-networks ofFDNN, that is,{F1,F 2,···,F |A|}, with non-sharing hidden layers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.036104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.908290Z digest=sha256:e79fb433d32dd6dfaf869f327c1b9d8c91d23405945fddaf470cbc3b7f327223

Observation 11e422ca-cc1a-43b4-9d64-cde127037e15 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.026020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.911478Z digest=sha256:e842ceaf17b4dfb80c243d549217ee0f974d4ea9b940841dfb80b0a106d1a262

Observation 40a15220-f7fa-43f0-886a-384e0e4f9b36 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.015633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.915042Z digest=sha256:011b46b8a57b9ccc616b3070232c258fb1e1e77b4f9c93208ed04a455d28894f

Observation 9e009a17-72c6-4c14-a707-f9738718d352 · outbound

This paper cites □ Different from the deep regression and classification problem, the convergence of the excess risk does not directly ensure the functional convergence of the estimated reward.

Learning Guarantee of Reward Modeling Using Deep Neural Networks □ Different from the deep regression and classification problem, the convergence of the excess risk does not directly ensure the functional convergence of the estimated reward

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.005167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.918108Z digest=sha256:a5a6b0dc78a14fed84989c11669cfcbc2362efa29049f8001c252d4a3e6681a2

Observation e5792b3a-01cc-4237-af5e-36dad5b3bfd8 · outbound

This paper cites and Bayati, M.

Learning Guarantee of Reward Modeling Using Deep Neural Networks and Bayati, M

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:53.875103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:53.875103Z digest=sha256:82c8332ceb2e6d0fd00d3bf7eab537f43a1f22dbd1b39e541c1937cfaaa1278a

Observation 9cf49d02-47d6-4dca-a9cc-20311bf1a2e1 · outbound

This paper cites Dual Active Learning for Reinforcement Learning from Human Feedback.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Dual Active Learning for Reinforcement Learning from Human Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:53.882277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:53.882277Z digest=sha256:bccdcff44f0f60b7182e7b9a0bb5d1f8ff61e69c2f36a806848c4e6a91ee500d

Observation 4a897fcd-bf4b-46a4-b36f-645274ee75c9 · outbound

This paper cites and Kupper, L.

Learning Guarantee of Reward Modeling Using Deep Neural Networks and Kupper, L

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.087748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.886562Z digest=sha256:8e7a3127b9c08cba8724436528fe9db4f5b03282e61e3614be2718dc021c36ec

Observation ff05ea61-5cf2-4448-bb54-cb05ad9c6d9f · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:53.893955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:53.893955Z digest=sha256:ec0b8c565eff02bc4667c1acaa45c30e94e4cc57f1b4b4c144162ff985484dc6

Observation 6dca5193-da27-4c43-a9f9-60b89abe5e22 · outbound

This paper cites B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M.

Learning Guarantee of Reward Modeling Using Deep Neural Networks B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M

Reference 1897

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.077909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:45:53.890107Z digest=sha256:fe945d6cb129eef993c91f42c939457a836ae8c08635c92009bb67160fea1876

Pith citing papers

No inbound Pith citation observations are available.