Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:14:45.567616Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2506.08712.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:14:45.567616Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:01:59.503415Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T14:02:39.828118Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1792dcec-7e4b-485d-aab6-2652c6154666 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Explaining individual predictions when features are dependent: More accurate approximations to shapley values
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98329f87-3bfb-457c-ba45-3f3c6e1eacf2 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Llama 3 model card
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 425b234e-dc88-4f52-914a-71c22079b1b3 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Mitigating reward over-optimization in direct alignment algorithms with adaptive importance sampling, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1638c1d0-5da9-4405-87c8-ac49348bd144 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization G., Guo, Z
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24dd7f57-28c4-4fe8-9074-5489596df715 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fceb882e-572d-44f2-beca-7305daa69a4c · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-level value preference optimization for mathematical reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeb05f53-2f33-4bc0-b3c9-4eb5bc98e9c0 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95b809b0-5d62-4004-b9d9-101bc6a66065 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Ultrafeedback: Boosting language models with high-quality feedback, 2023
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e194a62-4662-4361-ba68-22a6d3f5ea1d · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Enhancing chat language models by scaling high-quality instructional conversations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45141898-afe8-484c-8e62-c08716403b3e · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da680899-0219-4152-ad23-3e93ca90f203 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization KTO: Model Alignment as Prospect Theoretic Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31236e31-6e00-4b39-85ac-7f5232678988 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Scaling laws for reward model overoptimization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda1935b-ddae-4a4b-8f78-97fbe109079d · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Beyond imitation: Leveraging fine-grained quality signals for alignment
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a463a060-2d3c-49a0-b719-d5e49148b80e · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization A probabilistic earley parser as a psycholinguistic model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bfa7d60-0423-4c83-aebd-ddfe0126531f · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Reference-free monolithic preference optimization with odds ratio
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0350f2ed-d3ce-4d3c-830c-19bf1471817b · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Levy, R
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc52e0c3-8bb3-4d94-a507-dffcc102d1e6 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Mistral 7B
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d19fd706-3403-47a2-ad12-cb865ce56a1b · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b7bea3-d90f-4df8-87b0-84308a90de37 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization E., and Stoica, I
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e3df7f6-7a62-4f25-8a5d-58e625bfe14f · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebe34097-d2cb-4aaa-8b22-31ada1e6919a · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Not all tokens are what you need for pretraining
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 35ed17eb-b8c4-4dc6-a9fe-754b7d2fe017 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Provably mitigating overoptimization in RLHF : Your SFT loss is implicitly an adversarial regularizer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c436c405-a559-4406-a993-cd241f3d01e9 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Hutter, F
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1adf573f-560b-4750-9223-cce8d4ca61cf · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00841d81-e6b7-49be-92af-ad2c92ce077d · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Sim PO : Simple preference optimization with a reference-free reward
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0be6bca5-b5ca-4bc2-8d9a-e9b77595f50f · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization and Frank, S
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e0bd5a-75e2-468e-b903-a10285da3a45 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization F., Leike, J., and Lowe, R
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ee4273c-88b4-4a4d-9b2c-3d9c1bf7385f · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51e5e0f4-f6c6-48d5-9ce1-3c8eade8b949 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Disentangling Length from Quality in Direct Preference Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99509864-b239-4a70-b17e-fc59fcc97449 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization B., Finn, C., and Niekum, S
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97ba12ec-4d95-47d8-8976-42f0e7f567b6 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization D., Ermon, S., and Finn, C
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84fcbaca-6bd8-4148-bf89-5d8863154608 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58b40bde-26e0-4200-923a-e26b358609f2 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., and Martin, A
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cbce2f8-8b15-43d9-93b5-5f33853423d1 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5ae8f2-5af7-4907-8de6-8a392005d2f9 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Trl: Transformer reinforcement learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5991f4c3-c998-4191-811e-7aca3019706c · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization V., Murray, K., and Kim, Y
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ad94440-41a3-4cd0-9335-cea3803121d6 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Selective preference optimization via token-level reward function estimation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34656020-91f9-4927-a40e-e1992ba9a49a · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Hasegawa-Johnson, M
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d628485e-9912-4853-92ab-417c060d13ae · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Eom, S., Han, G., Nam, D., Jo, D., On, K.-W., Hasegawa-Johnson, M., Kim, S., and Yoo, C
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 683d2268-77fb-4ccf-95d4-895c8b07727f · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a63adae7-2697-41d5-b8ca-70c8752b2ccf · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5f835a-0a5b-41ae-ad1d-c0f68a006d57 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization S., Kim, J., and Yoo, C
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f05230-c189-4624-9aff-b930a294cd2c · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Tpc: Test-time procrustes calibration for diffusion-based human image animation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bba0a0e7-f868-48a1-bac1-32d45411538f · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization RRHF : Rank responses to align language models with human feedback
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1eeb1d21-fd0d-487f-a370-57350e86ea35 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization Token-level direct preference optimization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1dc57259-667c-4d60-b5b1-f7b8a9949c6b · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e796a6b-07b7-440a-873a-5c8207e6972e · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization P., Zhang, H., Gonzalez, J
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb49ef8-9c9b-4f85-9bb5-73958cbf26a9 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization T-REG: Preference Optimization with Token-Level Reward Regularization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0194a21-5383-4b36-b536-83db9bb9b377 · outbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization write newline
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ed9b45-e61e-4f44-81eb-d7947419bbf3 · inbound
Failure Modes of Maximum Entropy RLHF ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf7f5771-d720-4a92-87dc-47e4c6c247e1 · inbound
Normalized Rewards for Preference Optimization ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.