Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:05.082364Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2505.21395.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:05.082364Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 104 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 22f83a12-a945-4a69-908e-4ab81863aa5f · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a060b2f6-9f79-45b6-ad81-05cac294c2fc · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45695cbb-b815-4b7c-91bd-efdc0d7692df · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., and Sun, W
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2fea3a5-9512-4a06-9378-3726f05d2273 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Harnessing Density Ratios for Online Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e61b7ea-ceac-44e3-a682-6d065032228d · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Scalable Online Exploration via Coverability
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59df60b0-282d-45f0-881a-64d8913f321f · outbound
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b05808-b330-4592-ab14-99f21e382e4e · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., Schneider, J., and Ng, A
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63de237d-f678-4794-b06d-4da79c494ac9 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf6783ee-6407-423e-99ae-962abfa1c415 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Constitutional AI: Harmlessness from AI Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14710f35-8db2-49c9-93e0-44bb0f7bd651 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Contextual bandit algorithms with supervised learning guarantees
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a88ce9-3be9-4e59-bf16-9bcc8471aaa5 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 240e18bc-3414-4dcc-84a3-ded8531ff22f · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a629e6ad-7be2-4c03-81fd-d1473e5d3b6d · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294d0f4b-5166-4dcd-8e91-7a8a6a731517 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Dataset Reset Policy Optimization for RLHF
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c63e23e4-6167-4dce-9523-9a09316c548a · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Robust and private stochastic linear bandits
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51b7fda3-0a4e-485d-9cfa-2a913759b7d1 · outbound
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404f5f1f-107c-481a-9b74-dff23f57bfc9 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Jiang, N
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe4d398-1f28-4680-b100-e0cb89748d74 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sentenac, F
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d938f58b-be16-4a0d-843a-81517f30b9b4 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e50edb24-1c97-49c3-b789-41f0bbd070c8 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Distributed Differential Privacy in Multi-Armed Bandits
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264da0f3-a7cb-455b-94c4-2314d2555433 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Shuffle Private Linear Contextual Bandits
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e2ec5f-65c4-4eb4-b309-760c60063db9 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment R., Zhou, X., and Natarajan, N
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20de056-40a9-4f70-a0a4-d6ffa2a43c34 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34bfd27-8df7-4de2-aeec-e32b1e7132e0 · outbound
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e84e02-d205-473a-ac09-e48864c73dde · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Minimax-optimal off-policy evaluation with linear function approximation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe68c96-036c-4c61-ba86-8735a69f9b79 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Calibrating noise to sensitivity in private data analysis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3280134e-d360-4f8c-a9b0-dbcfe4d6476e · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8409717-f79f-4245-b9a6-2ea6febaafa3 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Importance-weighted offline learning done right
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c6f8f3-e850-427d-a039-f2b44a779a37 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment REBEL: Reinforcement Learning via Regressing Relative Rewards
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4675ef-62f4-4075-aedf-98bf12cc7089 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Local differential privacy for regret minimization in reinforcement learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc140ad7-f20b-4a3c-a01a-da6da7cd4447 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Hopkins, S
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12048df-d6ce-4d49-a913-cb2bba9b9466 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., Kamath, G., Majid, M., and Narayanan, S
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf75a51a-b4c8-4e9d-8880-63e46bd00e07 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b846a68-5427-4060-bbec-7f6dc4f75ef7 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcdcba6a-9b70-4803-abe7-3893d49db018 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8020fbf3-3226-454f-882d-00ff24c036a2 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.\ 5084--5096
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca989b1-0cf9-41af-b94d-c4c55d00eb5d · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Langford, J
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91464f4c-943b-4e25-b532-9f3ef1473a6d · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Broader Landscape of Robustness in Algorithmic Statistics
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14b5b11e-9475-4f5b-afa7-042b1927bccb · outbound
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 071cfa26-ec20-4869-b252-6081c14f389a · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Brown-Cohen, J
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad056ec6-d3dd-492f-8217-2ca786e0b09a · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Optidice: Offline policy optimization via stationary distribution correction estimation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9b1c00a-e13c-462e-b624-9bb8d9f8675d · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private linear bandits with partial distributed feedback
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2e13a3b-36b1-44a4-b8ba-79f9a6ba2927 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., and Yu, Y
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206d38f8-6001-4688-a32f-53524120eb29 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52856b81-da6d-4fc1-8ae8-d8eb8190a6cf · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Y., Yan, J., Jayaraman, D., and Bastani, O
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbd543d0-84ec-4ba1-afb1-dea967749296 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b5f535-eca7-4929-b022-dd2a8c11a880 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption Robust Offline Reinforcement Learning with Human Feedback
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1eb1efe-0e88-4f7e-b1e1-30674758ec0b · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Talwar, K
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57dbc156-7351-41c5-a40d-874bdce57ac9 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Thakurta, A
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d20f5511-e422-4b99-9a44-aa19f99bb3fd · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Szepesv \'a ri, C
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4825c11a-6490-491e-817b-18b2bfbc143e · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Nash Learning from Human Feedback
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f6a628-19e0-4ccb-8873-35e77c39a2fc · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b7c65ef-b82f-4a21-8d04-b3f32a471f12 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment ChatGPT : Optimizing language models for dialogue
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f5f7b64-9e18-40d0-b252-5700b8401841 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training language models to follow instructions with human feedback
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14c19cfa-deb2-487a-959f-8f20c231999f · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Wang, Y.-X
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9288cd28-495f-4282-b1f4-204b56685608 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., Ermon, S., and Finn, C
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d0142a0-ccc1-4823-9912-5ef0c8bdf33e · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bf8d535-f667-4f41-9793-1c298983eef5 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Multi-Armed Bandits with Local Differential Privacy
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4518b6ef-e776-4b72-8342-f7fcb33ae2f4 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Agnostic System Identification for Model-Based Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c612a59-241f-4f9b-a974-7f75e4c5f031 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc91668b-e6fc-4d75-9fa3-64bededec017 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08a18c92-2d24-44ba-aec1-fa6dbc4d9703 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Proximal Policy Optimization Algorithms
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39eca17c-cd02-4985-94a5-84b9be50371e · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 876b151e-b9d3-4fb6-92b3-4596e38486f6 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Benchmarks and Algorithms for Offline Preference-Based Reward Learning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7753a3a-af6e-4cb7-ab0c-504b6df18e02 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00b963f3-8f72-48a2-a7b2-787bae6547ce · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The importance of online data: Understanding preference fine-tuning via coverage
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a871c859-9728-4a2d-a6a5-96ddc928938a · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349350fa-2915-4818-9aba-b70be6b32109 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TrustLLM: Trustworthiness in Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee8bc4a-7444-4747-8fa6-c8e6159e49cf · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Principle-driven self-alignment of language models from scratch with minimal human supervision
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 528e43c5-8c35-4cc0-81c2-71cccdb87003 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment A Minimaximalist Approach to Reinforcement Learning from Human Feedback
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f00f59-19a3-4c2b-a2f4-8491679d2c76 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f84149dd-0a3d-42b8-8d5f-c638c3e517b2 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment LLaMA: Open and Efficient Foundation Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32fd6239-8a15-4a5e-8b43-1d4eb839d10e · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13631173-84fa-4065-84da-7aa8bb414245 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Private reinforcement learning with pac and regret guarantees
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb59928c-97bb-4b8e-80ad-1443f37636a1 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TRL : T ransformer R einforcement L earning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dcc33b2-c25b-4b26-9f24-290ac9944091 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Central Role of the Loss Function in Reinforcement Learning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 408a1957-edf2-4e6e-a5c0-60f311fc7d14 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Oracle-efficient pessimism: Offline policy optimization in contextual bandits
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc6e284f-0f30-4db7-b4a7-e9324e6af1ae · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is RLHF More Difficult than Standard RL?
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0913f4-fe91-46fc-8f3c-89eb5bc4d9c7 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d55713-5938-490b-a9f4-e55aa35bebe4 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4a93fb1-1819-4d37-9e85-7a6155634ec4 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Self-Play Preference Optimization for Language Model Alignment
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24c71555-cb6c-429e-81e9-b096df6f3489 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d05d8e8-2c3e-4b5a-81ae-0e73ee18c877 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c70c23d6-df86-41e0-8da4-2b2b588b97e4 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Foundations of Large Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f943956-9991-46a4-94aa-57ee8a6fb445 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman-consistent pessimism for offline reinforcement learning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47693abf-7fe6-4f98-9725-ed925a50e51a · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c102604-556e-4131-b27d-df8e3dfe5770 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Role of Coverage in Online Reinforcement Learning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24e11ad-5bc0-4e7a-a0df-1b77114804f8 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e623a67-03b9-4747-8d5a-b1b369fbcd82 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially Private Fine-tuning of Language Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c7d7ac-d4bb-42ab-90b3-8e8cca818102 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Offline reinforcement learning with realizability and single-policy concentrability
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a0bc272-c655-4f37-a4d1-9a8ef6274c35 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., and Sun, W
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13b583d9-506f-4e74-bd85-4e04f870217f · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption-robust offline reinforcement learning
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f25a33c-c908-4cf8-a2fd-d6b57a926c3f · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58fc734-8508-4caa-bbe5-520dddb3eb8d · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Locally differentially private (contextual) bandits learning
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b5a4910-4606-41c6-a9f4-f1cab496e3d7 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private reinforcement learning with linear function approximation
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8708e2fc-7628-4e8b-ad26-4173b24d4d05 · outbound
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6de8d3e-efe2-4eab-93c7-daf7f2c38802 · outbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Zhang, W
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.