Pith. sign in

Paper Citation Record · LEDGER

Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2404.01054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01054 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:08:39.194910Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T23:08:36.882178Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb7df4c2-e908-4bf5-9a24-4a4fb6cbb9b8 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.889530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:27c541d9dd2f07dbda286ca171df8471951b98d10b683bf4c9c14a5ba2d0975c

Observation 18beda7d-ee04-4a1c-814e-63e191cf3b4f · inbound

Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs cites this paper.

Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:03:54.337367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:03:54.337367Z digest=sha256:bce0d747e5e042cf9c9c0686d5dc278ce7e814901abe0b58b9fe35e0366399c0

Observation 4051c944-daed-4b13-a9a3-2cc5c1914d22 · inbound

Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation cites this paper.

Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:42:45.504260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:42:45.504260Z digest=sha256:cdd92bef1b51abb36dc7d97e7dd1d03c2172e02b3e9774c6db89229e82a8b2fa

Observation 78b88685-a8f4-4edc-bd62-785f86a1e35b · inbound

BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute cites this paper.

BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:05:27.667048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:05:27.667048Z digest=sha256:f6a2d7d8de8eeb5e80a5ca5f4a491db3544a891c03f7b7e2280b59ebc6e5c297

Observation 4a25a3e2-e9f6-4264-8b77-1bdda9d6d28e · inbound

AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning cites this paper.

AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:08:39.194910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:08:39.194910Z digest=sha256:f6f2e7341e1a4373c61a209bc608d8f69c06ba3ca6e84a9b34b69cc4eb0b7397

Observation 90838d63-31a2-46c7-8a5a-8ebd41642d03 · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.446650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.446650Z digest=sha256:2851742a5aaf7762b9e4a3d1067b3e83a58416a3c74bc568f4d3cbc8a0335a0a