Pith. sign in

Paper Citation Record · LEDGER

Unbiased Alignment for Large Language Models with Noisy Preferences

As of 6 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2607.03248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.03248 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T03:50:30.002103Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7685b021-cb95-4531-945e-73576eea0429 · outbound

This paper cites Qwen Technical Report.

Unbiased Alignment for Large Language Models with Noisy Preferences Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:41d5b31942a6f641996f01f38c27551135c944c3a4a85b7eaa587767d875265c

Observation 312ae219-22d4-4818-8bcc-956f263e0ab9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Unbiased Alignment for Large Language Models with Noisy Preferences Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:e977cb97ca40bf9e560c4b30801cf619acd5cb6d0173a090602631befa133da5

Observation 5c0c7a06-27ad-4677-982b-92c9c40f6844 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Unbiased Alignment for Large Language Models with Noisy Preferences D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:53ec0e1ec5113c3af173fbe79bf31f2e1378a4c457194b200d70ef6a591cfeb2

Observation 22df411c-19d2-46b3-ae7c-360413ce0342 · outbound

This paper cites The Llama 3 Herd of Models.

Unbiased Alignment for Large Language Models with Noisy Preferences The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:5fa388e25f7707d4a853d79a30526fc2bf1cb80f479124790e41193f3807afd8

Observation 472a3426-3be9-4046-b703-010870a5e3a7 · outbound

This paper cites H., Ghandeharioun, A., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R.

Unbiased Alignment for Large Language Models with Noisy Preferences H., Ghandeharioun, A., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:993c889b35975f6b179722c1130bf4209518abef7edd5c118be717f210ed247b

Observation 4575764d-2ac8-457d-b570-17abf0bea412 · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

Unbiased Alignment for Large Language Models with Noisy Preferences A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:012101a092cc152032bd5d8fcdae8d484ea3d19980fbab88b510266eb228e9c4

Observation f3da6bf0-28f8-4f54-92b3-d184b6060e54 · outbound

This paper cites DeepSeek-V3 Technical Report.

Unbiased Alignment for Large Language Models with Noisy Preferences DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:38ae52ec5bca1ad88bca6cb3a3f0dcdb6e08af517c2ddb0cdb3976ec810fe12e

Observation 5d92ddab-da28-442b-8778-3481b1b90723 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Unbiased Alignment for Large Language Models with Noisy Preferences WebGPT: Browser-assisted question-answering with human feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:ca0d50b893a3ce146e9576f86d4d0a03fc225a9e52f6db08d9ee977b9369da0a

Observation cffa97fa-0313-4b24-b93e-43b2cb96443b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Unbiased Alignment for Large Language Models with Noisy Preferences Proximal Policy Optimization Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:2ecc13371dccc76fbf2644dc396a1117b9e649ef232a4d16a1d99b4f82a4a15e

Observation 88dbfaf9-80ff-4b5d-b8ef-b564808f0418 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Unbiased Alignment for Large Language Models with Noisy Preferences DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:441b4546fed2f0699f5bf58e324d07bced659504f35c5d8cf8bfbbe7dd26bc75

Observation 4dd035b0-6c76-423b-94de-9fbaef639182 · outbound

This paper cites Qwen3 Technical Report.

Unbiased Alignment for Large Language Models with Noisy Preferences Qwen3 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:05dbd5b3d74aa29e8bc9be0681d834e0e5aefc3b5c8342c146f7ebd7335ea97e

Observation 1470d54a-8ab3-4f29-b8b4-1ffcaa06c8cf · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Unbiased Alignment for Large Language Models with Noisy Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:60cddd4b01e17c8e68b3c4cafd17ec2c7f5eeeda5e06184027bcbc746f2127e6

Observation c1f31224-f330-4bc4-904c-092c4b5b66f5 · outbound

This paper cites an unresolved cited work.

Unbiased Alignment for Large Language Models with Noisy Preferences Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:3397e0c68897b9dfa5564f2ed1e2f96120889e6cecfe62f1ec9892cd7edae5b7

Observation f847fc5d-9843-4ae7-a7bd-e7d97b8e041c · outbound

This paper cites an unresolved cited work.

Unbiased Alignment for Large Language Models with Noisy Preferences Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:e7a352576639a55f4f37addcf89a9adf16d0ecabdb7f8f4204a61ab32b3c81b9

Observation 77a050f8-c2a6-4c74-a8a0-476e51c04128 · outbound

This paper cites Proof of Corollary 4.10 Proof.

Unbiased Alignment for Large Language Models with Noisy Preferences Proof of Corollary 4.10 Proof

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:19a884e1ff67ff1337566920ea238e82c24dfa19c1f302da0289c2d5e52947a6

Observation 4a23be41-4f6f-4aa9-9774-ae2abc36f540 · outbound

This paper cites For reward model training, we train the model for 3 epochs with learning rate 1e-5.

Unbiased Alignment for Large Language Models with Noisy Preferences For reward model training, we train the model for 3 epochs with learning rate 1e-5

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:5d2cf89e191ce6ada8a9160324067edc059070d1a53643b15efd89f4c86a9e46

Pith citing papers

No inbound Pith citation observations are available.