Pith. sign in

Paper Citation Record · LEDGER

Quantile Regression for Distributional Reward Models in RLHF

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2409.10164.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.10164 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:40:34.367205Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:29.852309Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cafb5bb8-e090-46e9-96e6-715b3ae5428a · inbound

Data-adaptive Safety Rules for Training Reward Models cites this paper.

Data-adaptive Safety Rules for Training Reward Models Quantile Regression for Distributional Reward Models in RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:25:13.118759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:25:13.118759Z digest=sha256:813ce2561b92c4bfe579438b667d7a29b91f441afb3f46f308ace5552b4d1049

Observation 5d4c7c56-409a-4a5e-9726-ca7b20976493 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Quantile Regression for Distributional Reward Models in RLHF

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.819124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.819124Z digest=sha256:2073e94f29d73dd1f636bd34176da7b64360384522fdeb2ea039ccbade5fe8e3

Observation eebee092-3423-48c2-822b-92755e9743e4 · inbound

Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking cites this paper.

Sandcastles in the Storm: Revisiting the (Im)possibility of Strong Watermarking Quantile Regression for Distributional Reward Models in RLHF

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:40:34.367205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:40:34.367205Z digest=sha256:72c7991aa3e5d1a309a274534a49cbd73601675613e9a5a55653c7ccd215f6be

Observation 4ba916fa-e542-401b-acad-5514af589ed5 · inbound

Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning cites this paper.

Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning Quantile Regression for Distributional Reward Models in RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T22:24:50.863536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:24:50.863536Z digest=sha256:e6ec4c247659720d0b962e5860f42463ef933f27e3e4658e8b10522a2890ed28

Observation f331a955-9fea-44ab-9d3b-a9deb53005d7 · inbound

A Systematic Analysis of Base Model Choice for Reward Modeling cites this paper.

A Systematic Analysis of Base Model Choice for Reward Modeling Quantile Regression for Distributional Reward Models in RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:09:00.958423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:09:00.958423Z digest=sha256:6eef0ed242c37dfa2782ca050e3fda90780cf39602a04931dc75ef3754c3920a

Observation 2d18b654-3c6c-4188-805d-813beb713306 · inbound

Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge cites this paper.

Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge Quantile Regression for Distributional Reward Models in RLHF

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:53.890216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:53.890216Z digest=sha256:46dd2e464469afb05c7a331bb9a94290206696057c1486bd911f1e39be1f93be

Observation 53ffa0d5-a87a-4185-98c0-eff96bbce587 · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences Quantile Regression for Distributional Reward Models in RLHF

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:07.583666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:07.583666Z digest=sha256:f9d7d78769da9b9b1dc4199b5a10d3872b40a32dab8f546d6e2934081d4f3a56

Observation e14e7958-b5e2-40df-8bca-e24c582c732f · inbound

Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries cites this paper.

Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries Quantile Regression for Distributional Reward Models in RLHF

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:52.976745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:59:52.976745Z digest=sha256:a7e4491fa4f588823d83dee03177f29da96d55daa3897368ab784ad6c24aa9cd

Observation 605a7477-2042-4cab-a746-468a1fb2b8d0 · inbound

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization cites this paper.

PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization Quantile Regression for Distributional Reward Models in RLHF

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:44.017650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:42:44.017650Z digest=sha256:18864e2617c3005ec975f2223bc82aa7fc2a559af4c73edfcd2286ece5e1c678

Observation 33048014-125a-43ab-99e2-e3c720b8ab88 · inbound

Post-Training Large Language Models via Reinforcement Learning from Self-Feedback cites this paper.

Post-Training Large Language Models via Reinforcement Learning from Self-Feedback Quantile Regression for Distributional Reward Models in RLHF

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:53.895202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:17:53.895202Z digest=sha256:b9877c3cd89705f07fb924e35e60bae83d034734ea208dae9a0f79208109b3e6

Observation 1abfe77c-dd23-4c1d-94a2-7941aff43713 · inbound

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization cites this paper.

Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization Quantile Regression for Distributional Reward Models in RLHF

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:56.905829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:40:56.905829Z digest=sha256:81660ab8b7da08676825998a855e94a4e420cd9636e3a5f1a16c4cfa6de1790a

Observation ff967312-685a-45e0-8d41-ac7f87920a54 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Quantile Regression for Distributional Reward Models in RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:48:50.933078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:2983a17ae0ad75610db784476b3bf618312cd3e8f82b1739eba4c1e432a3ec4a

Observation 607041f9-2297-447a-8ecb-117f80fea90b · inbound

Hyperfastrl: Hypernetwork-based reinforcement learning for unified control of parametric chaotic PDEs cites this paper.

Hyperfastrl: Hypernetwork-based reinforcement learning for unified control of parametric chaotic PDEs Quantile Regression for Distributional Reward Models in RLHF

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:25:58.926309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:09:00.155657Z digest=sha256:60e56fab098042a1a769fe1d19746e7da27159dfa7f71d27e9e90307cde51e23

Observation 219033b7-07ce-4dd3-aef9-48b3bbbb5b42 · inbound

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context cites this paper.

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context Quantile Regression for Distributional Reward Models in RLHF

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.484521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T00:58:34.750132Z digest=sha256:38f2f2e87ce966d5938f92f4483812fd9a57ac7ab4201339b9391f73a18c98d9

Observation af1524c6-ab14-47d6-980b-fc1416bfd25c · inbound

StoryAlign: Evaluating and Training Reward Models for Story Generation cites this paper.

StoryAlign: Evaluating and Training Reward Models for Story Generation Quantile Regression for Distributional Reward Models in RLHF

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:31:07.636900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:29:13.549559Z digest=sha256:38e78b46487dc6d919eb9d70447b8712079e7552aaa623d0a006a132effecf79

Observation 51585fec-0444-44c3-baa0-d5dc12cf2a50 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF Quantile Regression for Distributional Reward Models in RLHF

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.853689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:e1c27d7d37e052493bc783d077f4e357de55975cd6abe413290cf03fcb90c0d8

Observation 2646d0b2-5b6e-49fb-a256-db1c4c2be5ad · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Quantile Regression for Distributional Reward Models in RLHF

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:25.589952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:25.589952Z digest=sha256:3fb10fbf87e7d711a8726d275a2de20dfad4f163cd91967518f21d772a865902