Pith. sign in

Paper Citation Record · LEDGER

Learning Guarantee of Reward Modeling Using Deep Neural Networks

As of 21 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2505.06601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06601 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:45:53.918108Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0329729c-1293-4c61-85e3-ccfa75fe2e25 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.107037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.870651Z digest=sha256:f26e139d53282ee3b950963f8e0513092afe7c2d43299f59bcb6d7edf3f12c4e

Observation 84cdfbd8-5826-4501-9843-5d8d24bbc0fa · outbound

This paper cites Now, we define two sets, given anyη∈(0,1), S1 ={s∈S: 0<r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)≤η}; S2 ={s∈S:r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)>η}.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Now, we define two sets, given anyη∈(0,1), S1 ={s∈S: 0<r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)≤η}; S2 ={s∈S:r ∗(s,πr∗(s))−max a∈A/πr∗(s) r∗ a(s)>η}

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.046769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.904919Z digest=sha256:55787cabeaa50f0eb984633d61608e693cb9d2ef315c614e4983bf8e2aedd780

Observation a370af60-ad00-4d31-a9cb-6e3fc2fe9c40 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.097126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.878712Z digest=sha256:b27a4b51505a9829af6a02645dea8f025cef0e222e93ba6aaafe9ebf11fbc6cf

Observation d2ec498a-836a-462a-8c4b-ca0dec5f7998 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.067826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.897920Z digest=sha256:0740b493afb499ac8b356e84de460b56c17008e09319f893a70454665b4a9a88

Observation daf2ee32-c232-43a1-bfa2-15b3c775353e · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.057379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.901469Z digest=sha256:b9c6d394459ecb72b062b4e9c9c2c7bbdafbb668632e8345aa1ac72fd4513757

Observation b0801597-7ea2-4518-b8f9-03d2e6f84c16 · outbound

This paper cites Without loss of generality, we also define classes of sub-networks ofFDNN, that is,{F1,F 2,···,F |A|}, with non-sharing hidden layers.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Without loss of generality, we also define classes of sub-networks ofFDNN, that is,{F1,F 2,···,F |A|}, with non-sharing hidden layers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.036104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.908290Z digest=sha256:ae42a431403a6319453ee289ab8cdc6fa5ea6a73d9ec854fbdaeb8734552c432

Observation 11e422ca-cc1a-43b4-9d64-cde127037e15 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.026020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.911478Z digest=sha256:60be00a478e49a06e76594d28a90029b04366d75cb05cb5ba925b962d0274109

Observation 40a15220-f7fa-43f0-886a-384e0e4f9b36 · outbound

This paper cites an unresolved cited work.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:45:54.015633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.915042Z digest=sha256:7df0c644c6d677b6b91ad307981e600014fb4e21899dda414f614a44456105cf

Observation 9e009a17-72c6-4c14-a707-f9738718d352 · outbound

This paper cites □ Different from the deep regression and classification problem, the convergence of the excess risk does not directly ensure the functional convergence of the estimated reward.

Learning Guarantee of Reward Modeling Using Deep Neural Networks □ Different from the deep regression and classification problem, the convergence of the excess risk does not directly ensure the functional convergence of the estimated reward

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.005167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.918108Z digest=sha256:29a7fffec8cc45a5ab7fddc794b79e2b203cf9cf7d527cc0535b9a7e74d8b07d

Observation e5792b3a-01cc-4237-af5e-36dad5b3bfd8 · outbound

This paper cites and Bayati, M.

Learning Guarantee of Reward Modeling Using Deep Neural Networks and Bayati, M

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:53.875103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:53.875103Z digest=sha256:8115cbe0ff53f07fe9b7b5f2caa9184e50c11fd7525c41b639246d24830afc40

Observation 9cf49d02-47d6-4dca-a9cc-20311bf1a2e1 · outbound

This paper cites Dual Active Learning for Reinforcement Learning from Human Feedback.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Dual Active Learning for Reinforcement Learning from Human Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:53.882277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:53.882277Z digest=sha256:0972574a378671337f3e92910f25bf2b1f46c1767844a5221e4d90241eae0bc4

Observation 4a897fcd-bf4b-46a4-b36f-645274ee75c9 · outbound

This paper cites and Kupper, L.

Learning Guarantee of Reward Modeling Using Deep Neural Networks and Kupper, L

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.087748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.886562Z digest=sha256:3198a608be20b9e95e20b3238d78f65ede105efba2819d1e2614099dd0d6d72b

Observation ff05ea61-5cf2-4448-bb54-cb05ad9c6d9f · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:53.893955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:53.893955Z digest=sha256:d6d9fa9b3e2deb7132de8fab89cd95351aa1c0b34c00d371e96a226e13a53dca

Observation 6dca5193-da27-4c43-a9f9-60b89abe5e22 · outbound

This paper cites B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M.

Learning Guarantee of Reward Modeling Using Deep Neural Networks B., Balakrishnan, S., Bradley, J., Parekh, A., Ramch, K., Wainwright, M

Reference 1897

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:54.077909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:53.890107Z digest=sha256:6c9fae5de0fc3e9852a0210aa2ce97a836e5a64e6631226ce4ea1fb1306a9fb4

Pith citing papers

No inbound Pith citation observations are available.