Pith. sign in

Paper Citation Record · LEDGER

Bayesian Reward Models for LLM Alignment

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2402.13210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.13210 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:44:44.755499Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:37:30.066691Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 66514782-85be-4daf-b14d-38c05902923e · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective Bayesian Reward Models for LLM Alignment

Reference 260

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.755499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.755499Z digest=sha256:60e4a5068d3dc388138ce6995655f75fd90a9e00cf4e16599f8675bd3291d9e3

Observation e343fe7e-4f12-44f1-84b3-6cf25f16acf5 · inbound

Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment cites this paper.

Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment Bayesian Reward Models for LLM Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:24:39.524257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:24:39.524257Z digest=sha256:55ea3f58dc0dee9fbb69c6dae276bd8d46a9f4d56e1c7a6a9323fcacb2b11257

Observation 7cf9db51-0e08-48f4-90c8-a2c8f098b12a · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Bayesian Reward Models for LLM Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.529784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.529784Z digest=sha256:6ed78307d88b5b398a132dcf238c0d255bf9a7b5d303af5b8d70d167a30c8abc

Observation 4440dea5-165d-4523-8693-b37bfd7be1bc · inbound

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives cites this paper.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian Reward Models for LLM Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.758976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.758976Z digest=sha256:7927a25544f8a14809edc13eab0a19217cb6ba0bfb47f3d70811cba40e42b998

Observation 80a23c92-a12d-4d19-8a64-564e89829535 · inbound

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning cites this paper.

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning Bayesian Reward Models for LLM Alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:37:30.068171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:07:06.227106Z digest=sha256:22cd4b2cfd2f331ddab7fc7551d9b75c045ab02b36e2f1fb449e71716d8be781

Observation fab0b4f4-7813-40ed-972f-2290c1d1aab7 · inbound

BaRA: Bayesian Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning cites this paper.

BaRA: Bayesian Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning Bayesian Reward Models for LLM Alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:28.489458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:55:13.502149Z digest=sha256:ac05d15f48739ff18d67c63df10de63de92e4b2887c180245324070fe1fa12fb

Observation 388f28e2-382c-41b3-b71c-84ee765bf8a4 · inbound

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI cites this paper.

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI Bayesian Reward Models for LLM Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T01:29:02.336552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:29:02.336552Z digest=sha256:2b2273f5f06eb67eca57aeda2d4bc7319c3a9ff5923334250afc46a33afb8508