Pith. sign in

Paper Citation Record · LEDGER

Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2402.19085.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19085 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:31:57.783767Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T21:12:58.860969Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 528954e4-c12f-4004-b160-bd3384c4d079 · inbound

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes cites this paper.

Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:42.606208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:38:42.606208Z digest=sha256:074413ca3cb0f46c529ad87131eea4ead2d663848889b29240ee98c1d25cf5ab

Observation d8dc0c14-1e56-425f-9c13-92f81bc530bc · inbound

e-SimFT: Alignment of Generative Models with Simulation Feedback for Pareto-Front Design Exploration cites this paper.

e-SimFT: Alignment of Generative Models with Simulation Feedback for Pareto-Front Design Exploration Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:11:21.939061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:11:21.939061Z digest=sha256:736f6f2f74387b88106362d85b1b04110cf900eb29c0b5518a240de3f5094e5d

Observation c1cbf768-ba07-4c76-a298-0511f59386b4 · inbound

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment cites this paper.

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:31:57.783767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:31:57.783767Z digest=sha256:d8557081aaf2668ebe6886f1cfdb067581846f030c905b2881735a3ffcf01bcb

Observation 2a0676ce-b45c-4989-bd44-a8b01b82fddc · inbound

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models cites this paper.

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:08:56.759061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:08:56.759061Z digest=sha256:621d12b52259c4694142aed575acf65804ce1c222268d14684f137a6d97c689b

Observation 99aba7be-d3b8-49e9-a7fd-c8ef588971ae · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:03.173700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:03.173700Z digest=sha256:f62d7c76fad09cff3a4434d1e605287b791b5f2e0ac93e08e2aa72cec751922c

Observation bafa858a-8dd3-4642-aa0c-05a8a4c77e12 · inbound

Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs cites this paper.

Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:13.529919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:13.529919Z digest=sha256:ac370329f4e368c8d96eaff96c75dbc96ff818af8bcd8258853939c3a4e1add4

Observation 5ed7d96a-2a3d-4258-bff5-fe7dc2be9f19 · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:43.394810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:43.394810Z digest=sha256:49e32f0bad4c6796cff394c7dcf16ff93298edb15407fa809705cc33c0003371

Observation 757d2aab-37d9-4c01-99b4-1280e0324221 · inbound

Learning to Control Summaries with Score Ranking cites this paper.

Learning to Control Summaries with Score Ranking Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.353139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T06:41:27.954392Z digest=sha256:f330bf7ad36fe5adc61b446317aab23838fca6ec952ede7de5b82a0c0dd191e0

Observation d32dcf92-9cce-4375-b1dc-d34cef5785b6 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.528098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:30ad1a2e0a24e4fefd8525c8b6454b32492caf8929f6e0e54fd3fd2ed11b43d1

Observation 05bb6824-eeca-49fa-b276-bfadedd7422c · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.862570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:caf106eb496d94b81554e250f24a74d44a6a5af504cde9f2afdb6c9d9b6d1cf7

Observation ef8a266c-270d-4d76-9b83-7f948be0c2f2 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 175

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:36.053105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:36.053105Z digest=sha256:3448e8b0a0173913ecc882b9224148764b211bb977aab6ebe17b1612b869c339

Observation 006d9a92-e3d2-4474-ae74-e729b55f641a · inbound

Procedural Fairness Failures in RLHF from Preference Averaging cites this paper.

Procedural Fairness Failures in RLHF from Preference Averaging Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.222855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.222855Z digest=sha256:46058f91dcd8a8be485fbc98b9e7f80c6cf57630fc0b519ed576a5a33b8aaf37