Pith. sign in

Paper Citation Record · LEDGER

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2507.07725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07725 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:38:10.381600Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 757181e0-3525-40f8-bb06-84aa17af40a3 · outbound

This paper cites In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.762128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.762128Z digest=sha256:bf4efcd1e83e1d52354beecaf311a35cfe2260841d368672db79e3bb70facdc6

Observation 96c56c30-3ddb-42ec-a8d2-3f4ce638959a · outbound

This paper cites the method of paired comparisons.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization the method of paired comparisons

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:38:10.807571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:38:09.793593Z digest=sha256:4a9b52354fc23c6d3fafcb6a82552138dafab9cd956a659f1c47ddc98de70995

Observation 2a3dfb02-53da-46da-a6b8-b9cb4713ac29 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.822571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.822571Z digest=sha256:6cbe4a941998eb30ed9f7da7bba4a8d6f6a90dda5c9f95f55dc29060939a1af9

Observation edf792c8-05a7-415b-8318-369e1b8a0856 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.880112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.880112Z digest=sha256:fd5123d91543e8a254c45e6895d1a4aeea2cb984ec24e4fc26c2a3be1c5645f0

Observation 9f0bc812-d474-4d61-bbfa-f4c4e9543a29 · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Direct Preference Knowledge Distillation for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:09.954900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:09.954900Z digest=sha256:5b23e7f14008990254be1415f34fda5f265f76716e4a1d68007c377f0bc5fb54

Observation 0f567fc3-240d-42c3-9de7-06984d22f8d7 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Rho-1: Not All Tokens Are What You Need

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.027182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.027182Z digest=sha256:f6e835ef1a6196197b4013d524c4557788bb38a2ee5a2269b0c03e99e73a7677

Observation b0468e90-10fe-4539-9ed8-a4cd3ff64abb · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.073796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.073796Z digest=sha256:15062be49ccae914641a5bf6dacc231aa2de9ace2603fadfcdc1ccd78a8b2970

Observation 3fd08862-22af-4123-86d5-2c1132d1fd50 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.166687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.166687Z digest=sha256:fb154912df24b7feeaf189cc5e1ce379fe6750ca3186c7a33c465085847eb57f

Observation 542a403c-c130-4b1b-b348-5626121a0de7 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.193064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.193064Z digest=sha256:3ea293a82da85e177ba92dd118bd561933a8136c8dcfff170b198714ceca750e

Observation 1e5aa232-9d66-42ed-861c-1a5db1eb57e1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.241141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.241141Z digest=sha256:2121def487c08ad84c18e713b93883b1f28a9f8f21ab3544c490b5fdf6468bd1

Observation e2a4ea7f-6a3c-4fa5-9f01-286a080192e0 · outbound

This paper cites an unresolved cited work.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization Unresolved cited work

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:38:10.624697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:38:10.308061Z digest=sha256:73764e24e59a213f1acbb8136a9ba02e31e9d15311d7ffdeaa35a1153e65226b

Observation 2aaae1ee-68f1-41d5-b2ac-a08da058766e · outbound

This paper cites , " * write output.state after.block = add.period write.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization , " * write output.state after.block = add.period write

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.339349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.339349Z digest=sha256:c59173050478a6cfb11a2850ac0207e92bfb6d5fc16abd7119aad82ed791ac3d

Observation 252f0a5b-4b4f-4d3e-a159-bab57d9c5aef · outbound

This paper cites write newline.

Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization write newline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:38:10.381600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:38:10.381600Z digest=sha256:c7b0c538146f47ba00ea90588961d8afc75fcaa7189dac7c4e3679c553affe88

Pith citing papers

No inbound Pith citation observations are available.