Pith. sign in

Paper Citation Record · LEDGER

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment

As of 10 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2505.21395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21395 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:05.082364Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 104 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22f83a12-a945-4a69-908e-4ab81863aa5f · outbound

This paper cites write newline.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.632371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.632371Z digest=sha256:1649baf6241318f8804d4f2665df1572b94c9efb40d69aa5bc7dec5685e8fec8

Observation a060b2f6-9f79-45b6-ad81-05cac294c2fc · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.693724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.693724Z digest=sha256:50d7d2f32d29a732af5b531556ea3d2f6e75551cde4bb9729c6aabaa3eaa4c98

Observation 45695cbb-b815-4b7c-91bd-efdc0d7692df · outbound

This paper cites M., and Sun, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., and Sun, W

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.808597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.808597Z digest=sha256:310c5b466055d04e67edea43614eb129115a5ac98b545a3d5c440596083c9825

Observation f2fea3a5-9512-4a06-9378-3726f05d2273 · outbound

This paper cites Harnessing Density Ratios for Online Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Harnessing Density Ratios for Online Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.889918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.889918Z digest=sha256:370169b5a9fd5fb9269f7925ce4c0ae089c465627f724c7707bf221dde18b30e

Observation 0e61b7ea-ceac-44e3-a682-6d065032228d · outbound

This paper cites Scalable Online Exploration via Coverability.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Scalable Online Exploration via Coverability

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:45:06.394885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:54.982699Z digest=sha256:69ce4d843e2c2118409bfdfab46d020aec7e469f11d4411f83ce79e587bc2ea7

Observation 59df60b0-282d-45f0-881a-64d8913f321f · outbound

This paper cites G., Guo, Z.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment G., Guo, Z

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.055754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.055754Z digest=sha256:704e05e08a9aaee165330af27e7048a9a471b6fa48dc722c453e548734db0a05

Observation 84b05808-b330-4592-ab14-99f21e382e4e · outbound

This paper cites M., Schneider, J., and Ng, A.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., Schneider, J., and Ng, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.169617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.169617Z digest=sha256:4be1081113c783737f2d615db103854bd2d7bff942a5c7ba950035f23a68a78b

Observation 63de237d-f678-4794-b06d-4da79c494ac9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.238757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.238757Z digest=sha256:940901f33dcebbbc5367dc476dfa5f923ef4bfbc2aee4ac5a799b11b20dc51b4

Observation cf6783ee-6407-423e-99ae-962abfa1c415 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.377088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.377088Z digest=sha256:2961e5d712b569d4085b3d49eee912520ef6d1e3394faa7791ec259e12ca7a7d

Observation 14710f35-8db2-49c9-93e0-44bb0f7bd651 · outbound

This paper cites Contextual bandit algorithms with supervised learning guarantees.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Contextual bandit algorithms with supervised learning guarantees

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.517503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.517503Z digest=sha256:ce92a05ad645543c54d979761fc75373e45459e7bf7f28e9f4dd6bcabee9bfdd

Observation 22a88ce9-3be9-4e59-bf16-9bcc8471aaa5 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.605096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.605096Z digest=sha256:6736df7a9a00d986f2afaf2fded7b6ac3be4b43bbaf200a32664bbf32d21e610

Observation 240e18bc-3414-4dcc-84a3-ded8531ff22f · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.718587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.718587Z digest=sha256:95777318afb5f9eff3e3b0cd8e3acbdf15549c0b1cb38c5af85c0e4ec3de0d0d

Observation a629e6ad-7be2-4c03-81fd-d1473e5d3b6d · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.826681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.826681Z digest=sha256:56e50223f32f4888415612313681038079eeacd6ea02438be9a78cc9c2e4c79d

Observation 294d0f4b-5166-4dcd-8e91-7a8a6a731517 · outbound

This paper cites Dataset Reset Policy Optimization for RLHF.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Dataset Reset Policy Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.917428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.917428Z digest=sha256:2031ab6d8d3dbe863bf21d8454585808a703a4b931bf09e6887740c0e2bc9ac6

Observation c63e23e4-6167-4dce-9523-9a09316c548a · outbound

This paper cites Robust and private stochastic linear bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Robust and private stochastic linear bandits

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.010458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.010458Z digest=sha256:74569add44ea3cc7e6ac48a1ad114d2082be78e7d5abbfbd1ab8e468b7baef9d

Observation 51b7fda3-0a4e-485d-9cfa-2a913759b7d1 · outbound

This paper cites and Hsu, D.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Hsu, D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.102692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.102692Z digest=sha256:015337a18448049cf47c9c49f74c94070212b10dee4682cd1408e8044b7482a5

Observation 404f5f1f-107c-481a-9b74-dff23f57bfc9 · outbound

This paper cites and Jiang, N.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Jiang, N

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.239470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.239470Z digest=sha256:2538e72088e1099909b24e36d297d1c0f196aaca11c8ae1be89235ceb83cc2bc

Observation 3fe4d398-1f28-4680-b100-e0cb89748d74 · outbound

This paper cites and Sentenac, F.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sentenac, F

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.337210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.337210Z digest=sha256:0f999e339288bda5c421696fdedda0606bb17d8d3c7f88d1a78e1c1cd6e8fce3

Observation d938f58b-be16-4a0d-843a-81517f30b9b4 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.414470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.414470Z digest=sha256:34ad0438407c2fd5bdb4e0b3e50dd40a0661ff05b2217ae1003d56b438f33f0f

Observation e50edb24-1c97-49c3-b789-41f0bbd070c8 · outbound

This paper cites Distributed Differential Privacy in Multi-Armed Bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Distributed Differential Privacy in Multi-Armed Bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.541880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.541880Z digest=sha256:ef72a32ba664e7ca6d2ffc12aa17e0fcce691d862110d05f47e4473b553d8608

Observation 264da0f3-a7cb-455b-94c4-2314d2555433 · outbound

This paper cites Shuffle Private Linear Contextual Bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Shuffle Private Linear Contextual Bandits

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.606840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.606840Z digest=sha256:a662bad0ce07e0eb8f83b3932400e6655c039f4b62fcb0f7a3a4e30756db7358

Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.673138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.673138Z digest=sha256:9b62489023e1c9480cb732f954f1d26b6aa9bf5ae5c67e1893974d31d27406aa

Observation 05e2ec5f-65c4-4eb4-b309-760c60063db9 · outbound

This paper cites R., Zhou, X., and Natarajan, N.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment R., Zhou, X., and Natarajan, N

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.734155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.734155Z digest=sha256:b3cad8af38a967945c6daf04dfe343578eda38d8330d5cc1d42ba88417ba3f61

Observation a20de056-40a9-4f70-a0a4-d6ffa2a43c34 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.793524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.793524Z digest=sha256:830b53183bae65f3199999b4a251ad9889d6c9abcf867c340e44675df119aa85

Observation d34bfd27-8df7-4de2-aeec-e32b1e7132e0 · outbound

This paper cites and Du, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Du, S

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.852370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.852370Z digest=sha256:f74b3771797540b8fa2db0e618288f89f963f62959458a3045f279bd78844e73

Observation 25e84e02-d205-473a-ac09-e48864c73dde · outbound

This paper cites Minimax-optimal off-policy evaluation with linear function approximation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Minimax-optimal off-policy evaluation with linear function approximation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.912450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.912450Z digest=sha256:1da4261392f995c3ec67487ab743206b96ab0eb569e97bd883375da972965cc6

Observation 2fe68c96-036c-4c61-ba86-8735a69f9b79 · outbound

This paper cites Calibrating noise to sensitivity in private data analysis.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Calibrating noise to sensitivity in private data analysis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.002442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.002442Z digest=sha256:4e058bb661d8a0fcf8147b1c1dc889ea1cc986a155f9ae2e5f045bb44acc46bb

Observation 3280134e-d360-4f8c-a9b0-dbcfe4d6476e · outbound

This paper cites Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.064925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.064925Z digest=sha256:cd9f562c4058e63b1bf5758ecaeb6b91e4c1382359d9945aefd59880ad145503

Observation b8409717-f79f-4245-b9a6-2ea6febaafa3 · outbound

This paper cites Importance-weighted offline learning done right.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Importance-weighted offline learning done right

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.137384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.137384Z digest=sha256:88eb367bfd18ac46e03b07c9831b65e0a286602487228810c6b6b0a6e276898d

Observation 47c6f8f3-e850-427d-a039-f2b44a779a37 · outbound

This paper cites REBEL: Reinforcement Learning via Regressing Relative Rewards.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment REBEL: Reinforcement Learning via Regressing Relative Rewards

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.195037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.195037Z digest=sha256:980525518922e86e6f960f3de5d37806522d33bcc36d9aa9ce988a56c36323bc

Observation eb4675ef-62f4-4075-aedf-98bf12cc7089 · outbound

This paper cites Local differential privacy for regret minimization in reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Local differential privacy for regret minimization in reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.243840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.243840Z digest=sha256:c462c6a5b2dea98a5144d570ec47a9f62cafc590975c5ee84beb1746f716d7ce

Observation bc140ad7-f20b-4a3c-a01a-da6da7cd4447 · outbound

This paper cites and Hopkins, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Hopkins, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.323106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.323106Z digest=sha256:b028378f2137f5d2d56cc8ec67f57ed0ae217c7145f4bcb1ef5c3e05bbe5203b

Observation f12048df-d6ce-4d49-a913-cb2bba9b9466 · outbound

This paper cites B., Kamath, G., Majid, M., and Narayanan, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., Kamath, G., Majid, M., and Narayanan, S

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.392906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.392906Z digest=sha256:cbf6a9dc1646b1e9ef676018b49c332922c807c1c2a79646442be0f24eb09e30

Observation cf75a51a-b4c8-4e9d-8880-63e46bd00e07 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.450115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.450115Z digest=sha256:e129947f8544539915a24c38004d71683b100c04904397574b58fcfd1bf3fa44

Observation 8b846a68-5427-4060-bbec-7f6dc4f75ef7 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.534624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.534624Z digest=sha256:3ac0c27ca879e32a8195874386fd9ec7cffaef469c3bc119b00b680861f8255b

Observation bcdcba6a-9b70-4803-abe7-3893d49db018 · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.616181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.616181Z digest=sha256:5dd0cc6e52456d275d77d6088da1e94501fefe54bcc989d9361c3efa7d0982bc

Observation 8020fbf3-3226-454f-882d-00ff24c036a2 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.\ 5084--5096.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.\ 5084--5096

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.676519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.676519Z digest=sha256:9aa9f49470aa6ed85b8e15f5ffeca1b5c46d6a15b717cd7edffc57c774688502

Observation 0ca989b1-0cf9-41af-b94d-c4c55d00eb5d · outbound

This paper cites and Langford, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Langford, J

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.751781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.751781Z digest=sha256:7d99c51f421d9850fc7c059d5c10f30e09a52ced7231b3964b95d2dc9f716fff

Observation 91464f4c-943b-4e25-b532-9f3ef1473a6d · outbound

This paper cites The Broader Landscape of Robustness in Algorithmic Statistics.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Broader Landscape of Robustness in Algorithmic Statistics

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:45:06.102730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.815194Z digest=sha256:e33cacab27c7f737ce835db97ab927d0fdf646b688ed6bca5545eaed1d158529

Observation 14b5b11e-9475-4f5b-afa7-042b1927bccb · outbound

This paper cites P., Lee, H.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment P., Lee, H

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.831953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.831953Z digest=sha256:da85823d67fbf2a45a4efb1397b12bc4d552012d237291fa116e690acaef98ad

Observation 071cfa26-ec20-4869-b252-6081c14f389a · outbound

This paper cites and Brown-Cohen, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Brown-Cohen, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.444969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.836692Z digest=sha256:58a08999d4684c6bd9e22093104082a66198a4b44c1fb17aa0c469912fd793ae

Observation ad056ec6-d3dd-492f-8217-2ca786e0b09a · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.208728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.840951Z digest=sha256:7364fa5e36bd61d2e26d1d62b5e80e5666f36c89841bada801d6c52c24693137

Observation d9b1c00a-e13c-462e-b624-9bb8d9f8675d · outbound

This paper cites Differentially private linear bandits with partial distributed feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private linear bandits with partial distributed feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.045189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.846124Z digest=sha256:0e3e5f3062e367c4d0ca7895f37e91e5d3badcae3a588a3253efd79853700845

Observation e2e13a3b-36b1-44a4-b8ba-79f9a6ba2927 · outbound

This paper cites B., and Yu, Y.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., and Yu, Y

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.887776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.862941Z digest=sha256:5643188283209b42827ac23ea73b5dc44bdcda9c64c8249f6570ca750d52cfe9

Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.881471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.881471Z digest=sha256:cdd6b2df434a79cb182bcf01e5d1fc82f5f5b113235bc009d86ba64c8042a632

Observation 206d38f8-6001-4688-a32f-53524120eb29 · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.905368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.905368Z digest=sha256:54a6f89d75e195da7fe4f9ef31eb379233f92045cbdfc24fb3bfca766007381f

Observation 52856b81-da6d-4fc1-8ae8-d8eb8190a6cf · outbound

This paper cites Y., Yan, J., Jayaraman, D., and Bastani, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Y., Yan, J., Jayaraman, D., and Bastani, O

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.640870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.963295Z digest=sha256:9c6bafa643b8e732a8f0bf3af7f20dc2394a840f2008427db50aeb7698049f5e

Observation bbd543d0-84ec-4ba1-afb1-dea967749296 · outbound

This paper cites Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.989120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.989120Z digest=sha256:536fafd9b64be0b0f41bdf6d8cf6a60ec6034b7cd32832ba0690087e779bdb29

Observation b7b5f535-eca7-4929-b022-dd2a8c11a880 · outbound

This paper cites Corruption Robust Offline Reinforcement Learning with Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption Robust Offline Reinforcement Learning with Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.018188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.018188Z digest=sha256:b78ef8448e0be37c38d40f81e4333a99f6f9f671a4c49ad021d2ae23c0a20544

Observation b1eb1efe-0e88-4f7e-b1e1-30674758ec0b · outbound

This paper cites and Talwar, K.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Talwar, K

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.067713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.067713Z digest=sha256:6b8b87c9b7c34d42706f342e622ede0b53da3365274d8873a6ed2229db335738

Observation 57dbc156-7351-41c5-a40d-874bdce57ac9 · outbound

This paper cites and Thakurta, A.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Thakurta, A

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.445447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.166200Z digest=sha256:68a38c0634aab95d14c13424d5d02683a71262355a9b218f543714876481d859

Observation d20f5511-e422-4b99-9a44-aa19f99bb3fd · outbound

This paper cites and Szepesv \'a ri, C.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Szepesv \'a ri, C

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.207803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.256381Z digest=sha256:951339f420092d5d6d81be8998235edc61dc18a5a33a83f3c5a2a9ec6e7fb02a

Observation 4825c11a-6490-491e-817b-18b2bfbc143e · outbound

This paper cites Nash Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Nash Learning from Human Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.339973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.339973Z digest=sha256:d845ff368f7b6b26c2ef0ac7322e25f0cae11a9bd2cf480bd4af3d75106b6ea0

Observation 07f6a628-19e0-4ccb-8873-35e77c39a2fc · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:45:12.005072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.441762Z digest=sha256:66b51db840b31965d59013c71bc063047bcba817b90df4405facf24558fdd139

Observation 1b7c65ef-b82f-4a21-8d04-b3f32a471f12 · outbound

This paper cites ChatGPT : Optimizing language models for dialogue.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment ChatGPT : Optimizing language models for dialogue

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.754842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.569884Z digest=sha256:5e2e805cb99dcf7dc0b5007180f6767ba3281e6ae4914a7b694031304338373a

Observation 6f5f7b64-9e18-40d0-b252-5700b8401841 · outbound

This paper cites Training language models to follow instructions with human feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training language models to follow instructions with human feedback

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.599409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.720589Z digest=sha256:f0e86898f8dfa172749ca1a0d4cdd2f8bdfe388bc05021e4b1d2fb85d79d28af

Observation 14c19cfa-deb2-487a-959f-8f20c231999f · outbound

This paper cites and Wang, Y.-X.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Wang, Y.-X

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.361226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.865665Z digest=sha256:6ce3b3e2ecb6ab16651cfe49edf37ab86cbba28de6b2bbf4f239c275c3670dbe

Observation 9288cd28-495f-4282-b1f4-204b56685608 · outbound

This paper cites D., Ermon, S., and Finn, C.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., Ermon, S., and Finn, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.164787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.994773Z digest=sha256:0829a4866e9f381c4ae7d6f45dc59e4e04b60e1b97d9ad9e9566935eb2eb9970

Observation 3d0142a0-ccc1-4823-9912-5ef0c8bdf33e · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.913677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:59.170713Z digest=sha256:620c076d144cc7a82c8c2590de939b21525b2fc5e10d0d99692ab4cf2e48e2ab

Observation 1bf8d535-f667-4f41-9793-1c298983eef5 · outbound

This paper cites Multi-Armed Bandits with Local Differential Privacy.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Multi-Armed Bandits with Local Differential Privacy

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.345604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.345604Z digest=sha256:4364156b613e7c7ea488ffb45a1c1500dcf0dfd41d47dbe83f8d37a29ec482ff

Observation 4518b6ef-e776-4b72-8342-f7fcb33ae2f4 · outbound

This paper cites Agnostic System Identification for Model-Based Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Agnostic System Identification for Model-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.521221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.521221Z digest=sha256:5b67505ddc347ab48fc72bcd5708ba2d0d7a3a6f5502fed494186f633976a21e

Observation 2c612a59-241f-4f9b-a974-7f75e4c5f031 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.736916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.736916Z digest=sha256:85fc461b42164b0ea1bd85a97614cc885ffc1bc6f646872bf4cd65e9d74b5b57

Observation cc91668b-e6fc-4d75-9fa3-64bededec017 · outbound

This paper cites and Sheffet, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.649260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:44:59.899253Z digest=sha256:eff2eea9b0011e9e65d094629068eabc3417fd0a3d75af9edacb0bf27292e282

Observation 08a18c92-2d24-44ba-aec1-fa6dbc4d9703 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Proximal Policy Optimization Algorithms

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.057955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.057955Z digest=sha256:71c5eea3be63d27c480bac6346fd04669ec0838ba7065c4fca6da9c463b22528

Observation 39eca17c-cd02-4985-94a5-84b9be50371e · outbound

This paper cites and Sheffet, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.438346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:00.151879Z digest=sha256:0ae94a68a57839a6eea9dd7e3b91cf6ac01c2137bb9f7fd2af8e112a42a2a825

Observation 876b151e-b9d3-4fb6-92b3-4596e38486f6 · outbound

This paper cites Benchmarks and Algorithms for Offline Preference-Based Reward Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.273913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.273913Z digest=sha256:86c93307a96a370110bc0e0e044a7c2bcbd2929af4670776833772bf4b9fdfe4

Observation e7753a3a-af6e-4cb7-ab0c-504b6df18e02 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.439564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.439564Z digest=sha256:74e3b40a3343a57ec82fe350414ed79c5b346581491b3920f4b179e0bf74d9f1

Observation 00b963f3-8f72-48a2-a7b2-787bae6547ce · outbound

This paper cites The importance of online data: Understanding preference fine-tuning via coverage.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The importance of online data: Understanding preference fine-tuning via coverage

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.250643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:00.549181Z digest=sha256:8f55e99680d8d88b97b17fd0e72578f741a143ac96530dfe5cd727ed398e6c91

Observation a871c859-9728-4a2d-a6a5-96ddc928938a · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.714395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.714395Z digest=sha256:f9376af6dcb7265d45f14678f247c956e168deb317939fdc298f38ff468635f5

Observation 349350fa-2915-4818-9aba-b70be6b32109 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TrustLLM: Trustworthiness in Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.876180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.876180Z digest=sha256:02078de1944b2bdb9f98a856c8f03f5859f0c3d61897171a889bec120a80195c

Observation cee8bc4a-7444-4747-8fa6-c8e6159e49cf · outbound

This paper cites Principle-driven self-alignment of language models from scratch with minimal human supervision.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Principle-driven self-alignment of language models from scratch with minimal human supervision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.086380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.019288Z digest=sha256:e29110e6e592ab21ddec14476287c609342aee03027d7d3551a24cfdcf7921e7

Observation 528e43c5-8c35-4cc0-81c2-71cccdb87003 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.173114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.173114Z digest=sha256:9e4c1e1d0ff1b00c7e1ead08f8a7e9248dce4435ab196676a79f19b4cd859d8a

Observation 31f00f59-19a3-4c2b-a2f4-8491679d2c76 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.361604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.361604Z digest=sha256:0b77d2f385bd2d8f571e242437f740ca683075e97dc598c33b8e636316abc3ef

Observation f84149dd-0a3d-42b8-8d5f-c638c3e517b2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.530685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.530685Z digest=sha256:f71c73a38ce28238b6fe473b9a145c7533999f3acaee6be3c629f871f82d5de5

Observation 32fd6239-8a15-4a5e-8b43-1d4eb839d10e · outbound

This paper cites Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.676167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.676167Z digest=sha256:251e7e8c220aa8b62892dd30538bea44f41172aa25ebb43bb01de2138599b672

Observation 13631173-84fa-4065-84da-7aa8bb414245 · outbound

This paper cites Private reinforcement learning with pac and regret guarantees.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Private reinforcement learning with pac and regret guarantees

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.913242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.801236Z digest=sha256:9ef38ac630fcdc9c5f74d79a5daa925ab0c656794a5322d48aec78f41bfb79cb

Observation eb59928c-97bb-4b8e-80ad-1443f37636a1 · outbound

This paper cites TRL : T ransformer R einforcement L earning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TRL : T ransformer R einforcement L earning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.749209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.936084Z digest=sha256:476b013e4b8efafaa15c6532206bbdb3884db447aaae50a7e67a4b32ee83d762

Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.062867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.062867Z digest=sha256:73321c71214fd75693d45a2bbc78afbee3b5508e1e7dd8d38f4ba28b0126dfdb

Observation 6dcc33b2-c25b-4b26-9f24-290ac9944091 · outbound

This paper cites The Central Role of the Loss Function in Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Central Role of the Loss Function in Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.217073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.217073Z digest=sha256:09d591195b250be20f33dcaa96a99cb0e82988428a0a7fb5b2e963f1e38210f6

Observation 408a1957-edf2-4e6e-a5c0-60f311fc7d14 · outbound

This paper cites Oracle-efficient pessimism: Offline policy optimization in contextual bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Oracle-efficient pessimism: Offline policy optimization in contextual bandits

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.514295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:02.369874Z digest=sha256:9392edaa6b9843d9970610a264eae59a6819da03b2a0a8fa572ebc6056c703a8

Observation bc6e284f-0f30-4db7-b4a7-e9324e6af1ae · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is RLHF More Difficult than Standard RL?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.561160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.561160Z digest=sha256:7287729b85c9d81f99b4ea435d0bf719ce759efabd81efa6cd0c181ef8dbf3d7

Observation ef0913f4-fe91-46fc-8f3c-89eb5bc4d9c7 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.755041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.755041Z digest=sha256:02044f400457b8f97d7f94b76d208e599bd616e95428e93a25626bbb177e62cc

Observation 38d55713-5938-490b-a9f4-e55aa35bebe4 · outbound

This paper cites On private and robust bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.353478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:02.884175Z digest=sha256:90a5a6a8580c12cd0b84fa10ea1ae7e98fd99817b48d8691fb8b10f4e9fe906a

Observation d4a93fb1-1819-4d37-9e85-7a6155634ec4 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Self-Play Preference Optimization for Language Model Alignment

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.046755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.046755Z digest=sha256:13f35c55fc22658d65dced869b19cd8781a1019fe38de4e060a223a779afddc6

Observation 24c71555-cb6c-429e-81e9-b096df6f3489 · outbound

This paper cites On private and robust bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.153167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.207476Z digest=sha256:2041eed2e19c3bce30272b328e5f838363361faebaae3dc2f068a0ee8ac15510

Observation 2d05d8e8-2c3e-4b5a-81ae-0e73ee18c877 · outbound

This paper cites On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.350482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.350482Z digest=sha256:0798e5e0d01453eed30e8480c2775830b9e32678e1e467b7e176310dee12d5bd

Observation c70c23d6-df86-41e0-8da4-2b2b588b97e4 · outbound

This paper cites Foundations of Large Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Foundations of Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.495496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.495496Z digest=sha256:3691f1902dae07ac8f3353c73135e933468d8bcf3536856fe12c5848c2a40f27

Observation 5f943956-9991-46a4-94aa-57ee8a6fb445 · outbound

This paper cites Bellman-consistent pessimism for offline reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman-consistent pessimism for offline reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.964046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.635728Z digest=sha256:c516ed0c9a879d696332e44950903fb912f4156d21e9976764ed357b8cac2942

Observation 47693abf-7fe6-4f98-9725-ed925a50e51a · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.783637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.708960Z digest=sha256:282a3d6c793d4a3b3b20fdabc0147e9c686412f858be1da9bf9f0e9fac502e60

Observation 2c102604-556e-4131-b27d-df8e3dfe5770 · outbound

This paper cites The Role of Coverage in Online Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Role of Coverage in Online Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.851174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.851174Z digest=sha256:91fbc9eb15e476ad8280a040dda26910e16ced93fb6c5ae73df10224385ba976

Observation f24e11ad-5bc0-4e7a-a0df-1b77114804f8 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.971759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.971759Z digest=sha256:a56903a2f4d65c5115bb232b27714cefaea5fb71b1c1a8c561d48c6eb9d2bf2f

Observation 1e623a67-03b9-4747-8d5a-b1b369fbcd82 · outbound

This paper cites Differentially Private Fine-tuning of Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially Private Fine-tuning of Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.056564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.056564Z digest=sha256:2f9a6d464dfa0822fd31820102b0bccda7d3cbb510388cf625b29be185921e0a

Observation 99c7d7ac-d4bb-42ab-90b3-8e8cca818102 · outbound

This paper cites Offline reinforcement learning with realizability and single-policy concentrability.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Offline reinforcement learning with realizability and single-policy concentrability

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.527175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.163239Z digest=sha256:9d09d59a6ccb1523b5fb112f69ff7c9b0beb527f63e3738a19a703498ad06a55

Observation 2a0bc272-c655-4f37-a4d1-9a8ef6274c35 · outbound

This paper cites D., and Sun, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., and Sun, W

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.315920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.365119Z digest=sha256:b25f3cc689ceb2fe9f390451903471a7ede31adb67bfa20b6209e40e1c954ada

Observation 13b583d9-506f-4e74-bd85-4e04f870217f · outbound

This paper cites Corruption-robust offline reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption-robust offline reinforcement learning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.096754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.456100Z digest=sha256:84aaed792c0f4456eafc1c8fddf036a2b387bbc1295139aa3f632145aeb49d1d

Observation 0f25a33c-c908-4cf8-a2fd-d6b57a926c3f · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.554366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.554366Z digest=sha256:37908fe2cbd211b81e4aaf08ff6d8c88e7aff89e4c670d490223a3304c62ce63

Observation b58fc734-8508-4caa-bbe5-520dddb3eb8d · outbound

This paper cites Locally differentially private (contextual) bandits learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Locally differentially private (contextual) bandits learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.834248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.710387Z digest=sha256:fb8c1e63997b66e7aa36af2ddc2621f676242a283429b6f15b8266e4b3126344

Observation 4b5a4910-4606-41c6-a9f4-f1cab496e3d7 · outbound

This paper cites Differentially private reinforcement learning with linear function approximation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private reinforcement learning with linear function approximation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.623841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.842225Z digest=sha256:fb0d974d6cc5c132fef24f58f165c98b08efcb79208c67704c745e8d78a8c24a

Observation 8708e2fc-7628-4e8b-ad26-4173b24d4d05 · outbound

This paper cites and Tan, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Tan, J

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.448453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.955804Z digest=sha256:137edd85b1f6369f1cbb0b0a4049672872aaaac7af4f2a8fde180cf930eebb88

Observation a6de8d3e-efe2-4eab-93c7-daf7f2c38802 · outbound

This paper cites and Zhang, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Zhang, W

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.202397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T13:45:05.082364Z digest=sha256:2fb9bbbe71b3b81ebcea71f722f028d6d53ede428865e6180adfbe3743c2a67d

Pith citing papers

No inbound Pith citation observations are available.