Pith. sign in

Paper Citation Record · LEDGER

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2506.08712.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08712 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:14:45.567616Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:01:59.503415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:02:39.828118Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1792dcec-7e4b-485d-aab6-2652c6154666 · outbound

This paper cites Explaining individual predictions when features are dependent: More accurate approximations to shapley values.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Explaining individual predictions when features are dependent: More accurate approximations to shapley values

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.444061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:41.625861Z digest=sha256:789035fbdb5542ca9c0ba62f427b5ba79dcd2bf99dcfd6c12285f554e51a0a67

Observation 98329f87-3bfb-457c-ba45-3f3c6e1eacf2 · outbound

This paper cites Llama 3 model card.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Llama 3 model card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:41.672739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:41.672739Z digest=sha256:46edd88f547c2d36d201baf594284aaf0d80a48b7dde937d04e970f4fc48b3d3

Observation 425b234e-dc88-4f52-914a-71c22079b1b3 · outbound

This paper cites Mitigating reward over-optimization in direct alignment algorithms with adaptive importance sampling, 2025.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Mitigating reward over-optimization in direct alignment algorithms with adaptive importance sampling, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.431146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:41.768753Z digest=sha256:579d40e831850767861c78591375d7f1e9a8093bf64d2147aebec24bd81e0a69

Observation 1638c1d0-5da9-4405-87c8-ac49348bd144 · outbound

This paper cites G., Guo, Z.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization G., Guo, Z

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.423282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:41.913775Z digest=sha256:afb1736fb573c552fe677c8cd3fc428404c1d50640f2b03c758eb9af444b344c

Observation 24dd7f57-28c4-4fe8-9074-5489596df715 · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.023438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.023438Z digest=sha256:8cfc2485b8f64b349421efbc3f737109dfea79c8de6bd9ddfce30c9bb6f77c1e

Observation fceb882e-572d-44f2-beca-7305daa69a4c · outbound

This paper cites Step-level value preference optimization for mathematical reasoning.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-level value preference optimization for mathematical reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.135215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.135215Z digest=sha256:c8f60aaf5189a8546d3e7decdc9d3ac0a655f8bdf2303be07047ce7dca593b31

Observation eeb05f53-2f33-4bc0-b3c9-4eb5bc98e9c0 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.415696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.210513Z digest=sha256:5d21b43633e6ec78ea5f0a73acf2a54830098b81c18238f3cbc5590db4f70835

Observation 95b809b0-5d62-4004-b9d9-101bc6a66065 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.408306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.296984Z digest=sha256:597043b4572b277750c6582d0e66500aae7be5225306bd4ed61ad53ca2220966

Observation 7e194a62-4662-4361-ba68-22a6d3f5ea1d · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Enhancing chat language models by scaling high-quality instructional conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.398724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.398724Z digest=sha256:22ed40c4310e2ad977d897cfae97157490333773f501835509a1674f9a764c65

Observation 45141898-afe8-484c-8e62-c08716403b3e · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.513413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.513413Z digest=sha256:12cb606d20ea7d459b990a579c77d71f03056f6790d7251f846c36fc0d709a9b

Observation da680899-0219-4152-ad23-3e93ca90f203 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.597895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.597895Z digest=sha256:a96b2192770cefc608a2d2255d409e763dc1f8b6eb455d9db00c25c893abd53a

Observation 31236e31-6e00-4b39-85ac-7f5232678988 · outbound

This paper cites Scaling laws for reward model overoptimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Scaling laws for reward model overoptimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.724379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.724379Z digest=sha256:f21ba22cbeb5893b74471d839ce362dab3de6847e277a8aa1cee98843fefda6f

Observation cda1935b-ddae-4a4b-8f78-97fbe109079d · outbound

This paper cites Beyond imitation: Leveraging fine-grained quality signals for alignment.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Beyond imitation: Leveraging fine-grained quality signals for alignment

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.395535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.807169Z digest=sha256:b088a1aab94fd838ef35ffc72d35e8355be31a9fd6491746f79635660694e832

Observation a463a060-2d3c-49a0-b719-d5e49148b80e · outbound

This paper cites A probabilistic earley parser as a psycholinguistic model.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization A probabilistic earley parser as a psycholinguistic model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.387379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:42.938224Z digest=sha256:86bbebde2f99f8d00d8a5c944b964345d46dd029166a9cbdbb08aa6e26e0e7ea

Observation 4bfa7d60-0423-4c83-aebd-ddfe0126531f · outbound

This paper cites Reference-free monolithic preference optimization with odds ratio.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Reference-free monolithic preference optimization with odds ratio

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.379614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.033275Z digest=sha256:b5d38d83955e0318cace5832a46f57446d3223116daa00a5282bca03eb2aeddb

Observation 0350f2ed-d3ce-4d3c-830c-19bf1471817b · outbound

This paper cites and Levy, R.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Levy, R

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.371360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.126624Z digest=sha256:eb9eb4d0322e23b6aeab134ff5086a2c6e53c42475a8d7f3245e57618ebfe8fe

Observation cc52e0c3-8bb3-4d94-a507-dffcc102d1e6 · outbound

This paper cites Mistral 7B.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Mistral 7B

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.201027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.201027Z digest=sha256:875a8d3c33ee52464221169ff158735808f9dbcc6d6c2b7d2ab26b4fe1b8ae1e

Observation d19fd706-3403-47a2-ad12-cb865ce56a1b · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.248348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.248348Z digest=sha256:90d0e0e2d24cb91355d886bf268facc2c85eb84c295a51640a3cd29c578f0c04

Observation 20b7bea3-d90f-4df8-87b0-84308a90de37 · outbound

This paper cites E., and Stoica, I.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization E., and Stoica, I

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.363481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.330957Z digest=sha256:68a91566137a6514078ec0c480b5bad90b63b458a2be23ca2f91307ab47ce725

Observation 4e3df7f6-7a62-4f25-8a5d-58e625bfe14f · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:47.355563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.419011Z digest=sha256:98589899e986ba947c53aa6d22eb79c3e0a512f446473543f63381e480763bc1

Observation ebe34097-d2cb-4aaa-8b22-31ada1e6919a · outbound

This paper cites Not all tokens are what you need for pretraining.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Not all tokens are what you need for pretraining

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.347856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.489684Z digest=sha256:489718167fabea10230124cbf5aa36de74b13c0d1ee30c58dcc5df378ff96695

Observation 35ed17eb-b8c4-4dc6-a9fe-754b7d2fe017 · outbound

This paper cites Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.340004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.553587Z digest=sha256:0f3215ba7ed79f75efec753d0a111f0d2fdd470137448b17d39c15a0fc09e862

Observation c436c405-a559-4406-a993-cd241f3d01e9 · outbound

This paper cites and Hutter, F.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Hutter, F

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.624256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.624256Z digest=sha256:cb4d93546497d90ef42ea75ee68e8fb4359e61278dc5c32261210764932d9ca6

Observation 1adf573f-560b-4750-9223-cce8d4ca61cf · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:47.326955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.698453Z digest=sha256:b34ec99d4bfe40092b4d71234dbb93bbc830d19113be1d6b9480b234ee7ba2af

Observation 00841d81-e6b7-49be-92af-ad2c92ce077d · outbound

This paper cites Sim PO : Simple preference optimization with a reference-free reward.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Sim PO : Simple preference optimization with a reference-free reward

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.318519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.768830Z digest=sha256:008e78cfaeb8d1a2223559b8472a42eeca828fe47d604e8c4f337561bbc272f6

Observation 0be6bca5-b5ca-4bc2-8d9a-e9b77595f50f · outbound

This paper cites and Frank, S.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Frank, S

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:43.851661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:43.851661Z digest=sha256:2e9497719f599041c90efe6aced8c66e9a8a9979d41da7e7a635c374daae3621

Observation c3e0bd5a-75e2-468e-b903-a10285da3a45 · outbound

This paper cites F., Leike, J., and Lowe, R.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization F., Leike, J., and Lowe, R

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.309553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:43.920632Z digest=sha256:8c556b63152114d5091e29d44528f579938d83aaefb88bfa86df70cdd365199e

Observation 2ee4273c-88b4-4a4d-9b2c-3d9c1bf7385f · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:47.262125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.027094Z digest=sha256:19a32cc0640140b7d09187b49c1af06e07b19a62ef0d47c47c51351e11b2c02c

Observation 51e5e0f4-f6c6-48d5-9ce1-3c8eade8b949 · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Disentangling Length from Quality in Direct Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.113845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.113845Z digest=sha256:8725c5fef58c34c52270f7a4086b174147ff07033da108d174ff7f19496d8b1b

Observation 99509864-b239-4a70-b17e-fc59fcc97449 · outbound

This paper cites B., Finn, C., and Niekum, S.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization B., Finn, C., and Niekum, S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:47.059754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.195859Z digest=sha256:f6b807d2aa5919154f28d85421d53221be64d7962fdaf1993b9be1aa673dfa47

Observation 97ba12ec-4d95-47d8-8976-42f0e7f567b6 · outbound

This paper cites D., Ermon, S., and Finn, C.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization D., Ermon, S., and Finn, C

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.766693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.285480Z digest=sha256:121d0523d139c1778bd76775cc8938f5985600d86376233f710671dc975655e5

Observation 84fcbaca-6bd8-4148-bf89-5d8863154608 · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:14:46.666687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.371677Z digest=sha256:cf07407e43f08f3b42ab787e74d77022bc50f4147753344feda7c58a26447d58

Observation 58b40bde-26e0-4200-923a-e26b358609f2 · outbound

This paper cites S., and Martin, A.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., and Martin, A

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.448597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.448597Z digest=sha256:438edd95a51438c9ffd80f57be7f905c253870c5fa8d2cdcddbfaf048d4357ec

Observation 4cbce2f8-8b15-43d9-93b5-5f33853423d1 · outbound

This paper cites an unresolved cited work.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.525429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.525429Z digest=sha256:f13715f4847d001c605afcb42cbbf6a0585c709660413fcf0bfa7c979748a94c

Observation db5ae8f2-5af7-4907-8de6-8a392005d2f9 · outbound

This paper cites Trl: Transformer reinforcement learning.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Trl: Transformer reinforcement learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.606004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.606004Z digest=sha256:514f3a3bca4f33180245edd7d065bd1fc5d1eadb4aa6505c2ac4fd00f4c40da2

Observation 5991f4c3-c998-4191-811e-7aca3019706c · outbound

This paper cites V., Murray, K., and Kim, Y.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization V., Murray, K., and Kim, Y

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.544726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.660895Z digest=sha256:35d45cced757e27ed0429198ea6e0b9a49b6a755f21321605a16a8f75d1fb6ee

Observation 8ad94440-41a3-4cd0-9335-cea3803121d6 · outbound

This paper cites Selective preference optimization via token-level reward function estimation.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Selective preference optimization via token-level reward function estimation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.737682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.737682Z digest=sha256:902b2e800c5e8680f3d1d88c0648d2a5c84d934c95e8029125c91a61d02fa9f9

Observation 34656020-91f9-4927-a40e-e1992ba9a49a · outbound

This paper cites S., Hasegawa-Johnson, M.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Hasegawa-Johnson, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.416795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:44.802381Z digest=sha256:68f90c877a360878b970f18560d6e7f3231f00ed2e8e0feb6ea7be8dfe5b8903

Observation d628485e-9912-4853-92ab-417c060d13ae · outbound

This paper cites S., Eom, S., Han, G., Nam, D., Jo, D., On, K.-W., Hasegawa-Johnson, M., Kim, S., and Yoo, C.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Eom, S., Han, G., Nam, D., Jo, D., On, K.-W., Hasegawa-Johnson, M., Kim, S., and Yoo, C

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.866870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.866870Z digest=sha256:63a272142242a57cf99032e3198abc8d25e780fd2214a4d79fa90eadc83dad0e

Observation 683d2268-77fb-4ccf-95d4-895c8b07727f · outbound

This paper cites ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.934080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:44.934080Z digest=sha256:7a5459e4ed9cea2f6b13f043ffdb0c4265e1e1ab7b3542014ceedc1586090351

Observation a63adae7-2697-41d5-b8ca-70c8752b2ccf · outbound

This paper cites C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.017306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.017306Z digest=sha256:6c9b8ceb5d7fa37c88e40e1f3c78be0457abce18e28d8b63478b101b3576b845

Observation af5f835a-0a5b-41ae-ad1d-c0f68a006d57 · outbound

This paper cites S., Kim, J., and Yoo, C.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Kim, J., and Yoo, C

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.080901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.080901Z digest=sha256:8cd0e36d1a9ea1cf02a8eedcb7db541cd32f69f79166a9b17c15f486ca57c797

Observation 77f05230-c189-4624-9aff-b930a294cd2c · outbound

This paper cites Tpc: Test-time procrustes calibration for diffusion-based human image animation.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Tpc: Test-time procrustes calibration for diffusion-based human image animation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.318669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.170101Z digest=sha256:71e8d4e19d6933eab8f7dced063b46543efc8d1dc0e524291aecf43011e582cc

Observation bba0a0e7-f868-48a1-bac1-32d45411538f · outbound

This paper cites RRHF : Rank responses to align language models with human feedback.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization RRHF : Rank responses to align language models with human feedback

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.199167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.244056Z digest=sha256:6e48167fb5add85da0699e4dab1fef7cbdf0693d5fdf9ee5731c42b2741873dd

Observation 1eeb1d21-fd0d-487f-a370-57350e86ea35 · outbound

This paper cites Token-level direct preference optimization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Token-level direct preference optimization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:14:46.081029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.277916Z digest=sha256:dd6e95ad545dc0a74e85d7c71e4b958351fb797b1bd1d5f98845fcb6597ae8ec

Observation 1dc57259-667c-4d60-b5b1-f7b8a9949c6b · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.342641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.342641Z digest=sha256:09a832ab2254dc6a9b234d7e198615a560e8880056898ae0fb7d770f8f0f331e

Observation 9e796a6b-07b7-440a-873a-5c8207e6972e · outbound

This paper cites P., Zhang, H., Gonzalez, J.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization P., Zhang, H., Gonzalez, J

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.422180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.422180Z digest=sha256:60d8024ff5e5131fad05707d1a9248e31ea07a785bfb6dd5d3df0885ff9a5c9f

Observation 2cb49ef8-9c9b-4f85-9bb5-73958cbf26a9 · outbound

This paper cites T-REG: Preference Optimization with Token-Level Reward Regularization.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization T-REG: Preference Optimization with Token-Level Reward Regularization

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:14:45.733451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:14:45.496001Z digest=sha256:93d4cac7048351384bf961f3d49fe72f7b92ec468676dc3f9f816909bf9ece65

Observation a0194a21-5383-4b36-b536-83db9bb9b377 · outbound

This paper cites write newline.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization write newline

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.567616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.567616Z digest=sha256:12f2e138feb3918959fa04aff1c3e94b469a2b47c1780e3200d9c6c1ad77f731

Pith citing papers

Observation 71ed9b45-e61e-4f44-81eb-d7947419bbf3 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.830504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:38817afc61d7b972a956476d222726e8cdf1f7a62fa073edef2b86485d69aefe

Observation bf7f5771-d720-4a92-87dc-47e4c6c247e1 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:59.503415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:59.503415Z digest=sha256:2eb2dd18c6a46101f1ba3cdf28e9426597a2c0951dcbcb7f64aea2230013e82f