Pith. sign in

Paper Citation Record · LEDGER

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment

As of 8 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2505.21395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21395 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:05.082364Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 104 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22f83a12-a945-4a69-908e-4ab81863aa5f · outbound

This paper cites write newline.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.632371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.632371Z digest=sha256:1649baf6241318f8804d4f2665df1572b94c9efb40d69aa5bc7dec5685e8fec8

Observation a060b2f6-9f79-45b6-ad81-05cac294c2fc · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.693724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.693724Z digest=sha256:50d7d2f32d29a732af5b531556ea3d2f6e75551cde4bb9729c6aabaa3eaa4c98

Observation 45695cbb-b815-4b7c-91bd-efdc0d7692df · outbound

This paper cites M., and Sun, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., and Sun, W

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.808597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.808597Z digest=sha256:310c5b466055d04e67edea43614eb129115a5ac98b545a3d5c440596083c9825

Observation f2fea3a5-9512-4a06-9378-3726f05d2273 · outbound

This paper cites Harnessing Density Ratios for Online Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Harnessing Density Ratios for Online Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:54.889918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:54.889918Z digest=sha256:1028feb41e5a84ef25508ce06dec85ee4084ad50e55557980636b6ed59fcda05

Observation 0e61b7ea-ceac-44e3-a682-6d065032228d · outbound

This paper cites Scalable Online Exploration via Coverability.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Scalable Online Exploration via Coverability

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:45:06.394885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:54.982699Z digest=sha256:304b3af492ad66ddcabee0d98c9a525f76b781bdbe6aecbf01a09b075cd00ea5

Observation 59df60b0-282d-45f0-881a-64d8913f321f · outbound

This paper cites G., Guo, Z.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment G., Guo, Z

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.055754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.055754Z digest=sha256:704e05e08a9aaee165330af27e7048a9a471b6fa48dc722c453e548734db0a05

Observation 84b05808-b330-4592-ab14-99f21e382e4e · outbound

This paper cites M., Schneider, J., and Ng, A.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment M., Schneider, J., and Ng, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.169617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.169617Z digest=sha256:4be1081113c783737f2d615db103854bd2d7bff942a5c7ba950035f23a68a78b

Observation 63de237d-f678-4794-b06d-4da79c494ac9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.238757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.238757Z digest=sha256:940901f33dcebbbc5367dc476dfa5f923ef4bfbc2aee4ac5a799b11b20dc51b4

Observation cf6783ee-6407-423e-99ae-962abfa1c415 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.377088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.377088Z digest=sha256:2961e5d712b569d4085b3d49eee912520ef6d1e3394faa7791ec259e12ca7a7d

Observation 14710f35-8db2-49c9-93e0-44bb0f7bd651 · outbound

This paper cites Contextual bandit algorithms with supervised learning guarantees.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Contextual bandit algorithms with supervised learning guarantees

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.517503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.517503Z digest=sha256:ce92a05ad645543c54d979761fc75373e45459e7bf7f28e9f4dd6bcabee9bfdd

Observation 22a88ce9-3be9-4e59-bf16-9bcc8471aaa5 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.605096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.605096Z digest=sha256:6736df7a9a00d986f2afaf2fded7b6ac3be4b43bbaf200a32664bbf32d21e610

Observation 240e18bc-3414-4dcc-84a3-ded8531ff22f · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.718587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.718587Z digest=sha256:95777318afb5f9eff3e3b0cd8e3acbdf15549c0b1cb38c5af85c0e4ec3de0d0d

Observation a629e6ad-7be2-4c03-81fd-d1473e5d3b6d · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.826681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.826681Z digest=sha256:56e50223f32f4888415612313681038079eeacd6ea02438be9a78cc9c2e4c79d

Observation 294d0f4b-5166-4dcd-8e91-7a8a6a731517 · outbound

This paper cites Dataset Reset Policy Optimization for RLHF.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Dataset Reset Policy Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:55.917428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:55.917428Z digest=sha256:2031ab6d8d3dbe863bf21d8454585808a703a4b931bf09e6887740c0e2bc9ac6

Observation c63e23e4-6167-4dce-9523-9a09316c548a · outbound

This paper cites Robust and private stochastic linear bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Robust and private stochastic linear bandits

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.010458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.010458Z digest=sha256:74569add44ea3cc7e6ac48a1ad114d2082be78e7d5abbfbd1ab8e468b7baef9d

Observation 51b7fda3-0a4e-485d-9cfa-2a913759b7d1 · outbound

This paper cites and Hsu, D.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Hsu, D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.102692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.102692Z digest=sha256:015337a18448049cf47c9c49f74c94070212b10dee4682cd1408e8044b7482a5

Observation 404f5f1f-107c-481a-9b74-dff23f57bfc9 · outbound

This paper cites and Jiang, N.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Jiang, N

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.239470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.239470Z digest=sha256:2538e72088e1099909b24e36d297d1c0f196aaca11c8ae1be89235ceb83cc2bc

Observation 3fe4d398-1f28-4680-b100-e0cb89748d74 · outbound

This paper cites and Sentenac, F.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sentenac, F

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.337210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.337210Z digest=sha256:0f999e339288bda5c421696fdedda0606bb17d8d3c7f88d1a78e1c1cd6e8fce3

Observation d938f58b-be16-4a0d-843a-81517f30b9b4 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.414470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.414470Z digest=sha256:34ad0438407c2fd5bdb4e0b3e50dd40a0661ff05b2217ae1003d56b438f33f0f

Observation e50edb24-1c97-49c3-b789-41f0bbd070c8 · outbound

This paper cites Distributed Differential Privacy in Multi-Armed Bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Distributed Differential Privacy in Multi-Armed Bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.541880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.541880Z digest=sha256:ef72a32ba664e7ca6d2ffc12aa17e0fcce691d862110d05f47e4473b553d8608

Observation 264da0f3-a7cb-455b-94c4-2314d2555433 · outbound

This paper cites Shuffle Private Linear Contextual Bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Shuffle Private Linear Contextual Bandits

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.606840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.606840Z digest=sha256:17035754bac01bd722d964e2594932cc9bd02af1d9ca555b59f121138a8c73e9

Observation c07e7b14-0763-4b65-88e6-5d97b0cb27d7 · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.673138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.673138Z digest=sha256:1730db6e1eebc87514dcfc52ae1c5d3086a354c2323429da79c32f0c7a637c8a

Observation 05e2ec5f-65c4-4eb4-b309-760c60063db9 · outbound

This paper cites R., Zhou, X., and Natarajan, N.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment R., Zhou, X., and Natarajan, N

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.734155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.734155Z digest=sha256:b3cad8af38a967945c6daf04dfe343578eda38d8330d5cc1d42ba88417ba3f61

Observation a20de056-40a9-4f70-a0a4-d6ffa2a43c34 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.793524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.793524Z digest=sha256:830b53183bae65f3199999b4a251ad9889d6c9abcf867c340e44675df119aa85

Observation d34bfd27-8df7-4de2-aeec-e32b1e7132e0 · outbound

This paper cites and Du, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Du, S

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.852370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.852370Z digest=sha256:f74b3771797540b8fa2db0e618288f89f963f62959458a3045f279bd78844e73

Observation 25e84e02-d205-473a-ac09-e48864c73dde · outbound

This paper cites Minimax-optimal off-policy evaluation with linear function approximation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Minimax-optimal off-policy evaluation with linear function approximation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:56.912450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:56.912450Z digest=sha256:1da4261392f995c3ec67487ab743206b96ab0eb569e97bd883375da972965cc6

Observation 2fe68c96-036c-4c61-ba86-8735a69f9b79 · outbound

This paper cites Calibrating noise to sensitivity in private data analysis.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Calibrating noise to sensitivity in private data analysis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.002442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.002442Z digest=sha256:4e058bb661d8a0fcf8147b1c1dc889ea1cc986a155f9ae2e5f045bb44acc46bb

Observation 3280134e-d360-4f8c-a9b0-dbcfe4d6476e · outbound

This paper cites Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exposing Privacy Gaps: Membership Inference Attack on Preference Data for LLM Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.064925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.064925Z digest=sha256:cd9f562c4058e63b1bf5758ecaeb6b91e4c1382359d9945aefd59880ad145503

Observation b8409717-f79f-4245-b9a6-2ea6febaafa3 · outbound

This paper cites Importance-weighted offline learning done right.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Importance-weighted offline learning done right

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.137384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.137384Z digest=sha256:88eb367bfd18ac46e03b07c9831b65e0a286602487228810c6b6b0a6e276898d

Observation 47c6f8f3-e850-427d-a039-f2b44a779a37 · outbound

This paper cites REBEL: Reinforcement Learning via Regressing Relative Rewards.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment REBEL: Reinforcement Learning via Regressing Relative Rewards

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.195037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.195037Z digest=sha256:980525518922e86e6f960f3de5d37806522d33bcc36d9aa9ce988a56c36323bc

Observation eb4675ef-62f4-4075-aedf-98bf12cc7089 · outbound

This paper cites Local differential privacy for regret minimization in reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Local differential privacy for regret minimization in reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.243840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.243840Z digest=sha256:c462c6a5b2dea98a5144d570ec47a9f62cafc590975c5ee84beb1746f716d7ce

Observation bc140ad7-f20b-4a3c-a01a-da6da7cd4447 · outbound

This paper cites and Hopkins, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Hopkins, S

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.323106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.323106Z digest=sha256:b028378f2137f5d2d56cc8ec67f57ed0ae217c7145f4bcb1ef5c3e05bbe5203b

Observation f12048df-d6ce-4d49-a913-cb2bba9b9466 · outbound

This paper cites B., Kamath, G., Majid, M., and Narayanan, S.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., Kamath, G., Majid, M., and Narayanan, S

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.392906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.392906Z digest=sha256:cbf6a9dc1646b1e9ef676018b49c332922c807c1c2a79646442be0f24eb09e30

Observation cf75a51a-b4c8-4e9d-8880-63e46bd00e07 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.450115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.450115Z digest=sha256:e129947f8544539915a24c38004d71683b100c04904397574b58fcfd1bf3fa44

Observation 8b846a68-5427-4060-bbec-7f6dc4f75ef7 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.534624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.534624Z digest=sha256:3ac0c27ca879e32a8195874386fd9ec7cffaef469c3bc119b00b680861f8255b

Observation bcdcba6a-9b70-4803-abe7-3893d49db018 · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.616181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.616181Z digest=sha256:5dd0cc6e52456d275d77d6088da1e94501fefe54bcc989d9361c3efa7d0982bc

Observation 8020fbf3-3226-454f-882d-00ff24c036a2 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.\ 5084--5096.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pp.\ 5084--5096

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.676519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.676519Z digest=sha256:9aa9f49470aa6ed85b8e15f5ffeca1b5c46d6a15b717cd7edffc57c774688502

Observation 0ca989b1-0cf9-41af-b94d-c4c55d00eb5d · outbound

This paper cites and Langford, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Langford, J

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.751781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.751781Z digest=sha256:7d99c51f421d9850fc7c059d5c10f30e09a52ced7231b3964b95d2dc9f716fff

Observation 91464f4c-943b-4e25-b532-9f3ef1473a6d · outbound

This paper cites The Broader Landscape of Robustness in Algorithmic Statistics.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Broader Landscape of Robustness in Algorithmic Statistics

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:45:06.102730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.815194Z digest=sha256:f4eac0f1bae1ed7c6a42e66769a13d5f12e0884a1033a9935a09c7b568eebb24

Observation 14b5b11e-9475-4f5b-afa7-042b1927bccb · outbound

This paper cites P., Lee, H.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment P., Lee, H

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.831953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.831953Z digest=sha256:da85823d67fbf2a45a4efb1397b12bc4d552012d237291fa116e690acaef98ad

Observation 071cfa26-ec20-4869-b252-6081c14f389a · outbound

This paper cites and Brown-Cohen, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Brown-Cohen, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.444969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.836692Z digest=sha256:a19734eb674929fd092ccb55ea056d6a94a08faef3a935b4cc28ca1bfb5f0aac

Observation ad056ec6-d3dd-492f-8217-2ca786e0b09a · outbound

This paper cites Optidice: Offline policy optimization via stationary distribution correction estimation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Optidice: Offline policy optimization via stationary distribution correction estimation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.208728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.840951Z digest=sha256:2d154156182f76c38a8732be080bcbac282a2ab62b771a1499fc4a97fae09584

Observation d9b1c00a-e13c-462e-b624-9bb8d9f8675d · outbound

This paper cites Differentially private linear bandits with partial distributed feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private linear bandits with partial distributed feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:13.045189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.846124Z digest=sha256:f1ba0614a612c1da7948280676549469a2ebd4534c0cfe0ad479ef41f86c0eb4

Observation e2e13a3b-36b1-44a4-b8ba-79f9a6ba2927 · outbound

This paper cites B., and Yu, Y.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment B., and Yu, Y

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.887776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.862941Z digest=sha256:e44e974abec4a3565deb33d2f20d57df9d9692d73611ec6669c64c7369837ddc

Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.881471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.881471Z digest=sha256:abf25b1906ad4124a1bdcbd80068b99a27775bf2e5d1333532070d2fc6e0f464

Observation 206d38f8-6001-4688-a32f-53524120eb29 · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.905368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.905368Z digest=sha256:54a6f89d75e195da7fe4f9ef31eb379233f92045cbdfc24fb3bfca766007381f

Observation 52856b81-da6d-4fc1-8ae8-d8eb8190a6cf · outbound

This paper cites Y., Yan, J., Jayaraman, D., and Bastani, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Y., Yan, J., Jayaraman, D., and Bastani, O

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.640870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:57.963295Z digest=sha256:5e98e009c9a587d2a0e45d819d74b5ec918c24cb05db629ff078bd431d7fb0de

Observation bbd543d0-84ec-4ba1-afb1-dea967749296 · outbound

This paper cites Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.989120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.989120Z digest=sha256:536fafd9b64be0b0f41bdf6d8cf6a60ec6034b7cd32832ba0690087e779bdb29

Observation b7b5f535-eca7-4929-b022-dd2a8c11a880 · outbound

This paper cites Corruption Robust Offline Reinforcement Learning with Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption Robust Offline Reinforcement Learning with Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.018188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.018188Z digest=sha256:97db1cdd3d939784affae731f7cbf07f69fb107815de94948d93f3dee4bcb269

Observation b1eb1efe-0e88-4f7e-b1e1-30674758ec0b · outbound

This paper cites and Talwar, K.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Talwar, K

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.067713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.067713Z digest=sha256:6b8b87c9b7c34d42706f342e622ede0b53da3365274d8873a6ed2229db335738

Observation 57dbc156-7351-41c5-a40d-874bdce57ac9 · outbound

This paper cites and Thakurta, A.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Thakurta, A

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.445447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.166200Z digest=sha256:1ae52b76e65f350ca7b8cf6100e1bc5b286292c0562765ec68b34c88acbf9d89

Observation d20f5511-e422-4b99-9a44-aa19f99bb3fd · outbound

This paper cites and Szepesv \'a ri, C.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Szepesv \'a ri, C

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:12.207803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.256381Z digest=sha256:0da2f8aa2f24bb577f883fa7b807f5cffc50f7d38ef53982c1500544e62ae753

Observation 4825c11a-6490-491e-817b-18b2bfbc143e · outbound

This paper cites Nash Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Nash Learning from Human Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:58.339973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:58.339973Z digest=sha256:d845ff368f7b6b26c2ef0ac7322e25f0cae11a9bd2cf480bd4af3d75106b6ea0

Observation 07f6a628-19e0-4ccb-8873-35e77c39a2fc · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:45:12.005072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.441762Z digest=sha256:4c4b44c3c9a532db09989efc2cfcef1f16a027f4d811e277ee4f41c33532355c

Observation 1b7c65ef-b82f-4a21-8d04-b3f32a471f12 · outbound

This paper cites ChatGPT : Optimizing language models for dialogue.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment ChatGPT : Optimizing language models for dialogue

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.754842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.569884Z digest=sha256:f873dc3b3a78b534f5258f2eca3662b30cd39446bd450138cb470132fc4ef60c

Observation 6f5f7b64-9e18-40d0-b252-5700b8401841 · outbound

This paper cites Training language models to follow instructions with human feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Training language models to follow instructions with human feedback

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.599409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.720589Z digest=sha256:ffa0fa4315180c95c676d311f01059259e68256c9a3bc47e941e53b8b42e1475

Observation 14c19cfa-deb2-487a-959f-8f20c231999f · outbound

This paper cites and Wang, Y.-X.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Wang, Y.-X

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.361226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.865665Z digest=sha256:03658364cdf529d8c44f027b2ac4168f42d54ebb2391bcb39bfcf32951c50c9a

Observation 9288cd28-495f-4282-b1f4-204b56685608 · outbound

This paper cites D., Ermon, S., and Finn, C.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., Ermon, S., and Finn, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:11.164787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:58.994773Z digest=sha256:300dde6886347ae0d3236a8caed775b3d5636b66b2d76a41f4b84613b64633ef

Observation 3d0142a0-ccc1-4823-9912-5ef0c8bdf33e · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.913677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:59.170713Z digest=sha256:619222e3bf486e13d0719f119229f6ac8adb2c2c9325bb695c7724c4af82060a

Observation 1bf8d535-f667-4f41-9793-1c298983eef5 · outbound

This paper cites Multi-Armed Bandits with Local Differential Privacy.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Multi-Armed Bandits with Local Differential Privacy

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.345604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.345604Z digest=sha256:38bb897ed14f791deafdacf449dd496a89fd4acda609b11392e1d386c5e71767

Observation 4518b6ef-e776-4b72-8342-f7fcb33ae2f4 · outbound

This paper cites Agnostic System Identification for Model-Based Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Agnostic System Identification for Model-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.521221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.521221Z digest=sha256:5b67505ddc347ab48fc72bcd5708ba2d0d7a3a6f5502fed494186f633976a21e

Observation 2c612a59-241f-4f9b-a974-7f75e4c5f031 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:59.736916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:59.736916Z digest=sha256:85fc461b42164b0ea1bd85a97614cc885ffc1bc6f646872bf4cd65e9d74b5b57

Observation cc91668b-e6fc-4d75-9fa3-64bededec017 · outbound

This paper cites and Sheffet, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.649260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:44:59.899253Z digest=sha256:86dd7ee0d9217fa2e9b90076c8641d7e45b9c5e186f317811b47d6b0ec437644

Observation 08a18c92-2d24-44ba-aec1-fa6dbc4d9703 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Proximal Policy Optimization Algorithms

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.057955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.057955Z digest=sha256:71c5eea3be63d27c480bac6346fd04669ec0838ba7065c4fca6da9c463b22528

Observation 39eca17c-cd02-4985-94a5-84b9be50371e · outbound

This paper cites and Sheffet, O.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Sheffet, O

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.438346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:00.151879Z digest=sha256:0a4abdc7a1f9c1fb22fe66a713531d836a0649efa62b8e206ba4cf833297f087

Observation 876b151e-b9d3-4fb6-92b3-4596e38486f6 · outbound

This paper cites Benchmarks and Algorithms for Offline Preference-Based Reward Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.273913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.273913Z digest=sha256:b627ad7118b2ffcf0e8271571e2096de9490a09c1520092d507d85a336c36f94

Observation e7753a3a-af6e-4cb7-ab0c-504b6df18e02 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.439564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.439564Z digest=sha256:74e3b40a3343a57ec82fe350414ed79c5b346581491b3920f4b179e0bf74d9f1

Observation 00b963f3-8f72-48a2-a7b2-787bae6547ce · outbound

This paper cites The importance of online data: Understanding preference fine-tuning via coverage.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The importance of online data: Understanding preference fine-tuning via coverage

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.250643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:00.549181Z digest=sha256:18a95c2623a6955d016d4c42ed650fa074d8eb58e2538f5e078d661bc03f3e2b

Observation a871c859-9728-4a2d-a6a5-96ddc928938a · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.714395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.714395Z digest=sha256:f9376af6dcb7265d45f14678f247c956e168deb317939fdc298f38ff468635f5

Observation 349350fa-2915-4818-9aba-b70be6b32109 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TrustLLM: Trustworthiness in Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.876180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.876180Z digest=sha256:02078de1944b2bdb9f98a856c8f03f5859f0c3d61897171a889bec120a80195c

Observation cee8bc4a-7444-4747-8fa6-c8e6159e49cf · outbound

This paper cites Principle-driven self-alignment of language models from scratch with minimal human supervision.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Principle-driven self-alignment of language models from scratch with minimal human supervision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:10.086380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.019288Z digest=sha256:ed66c2d36fe966df36b044c12c38b4210eb4621968635bd8792bf96745daef19

Observation 528e43c5-8c35-4cc0-81c2-71cccdb87003 · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.173114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.173114Z digest=sha256:2cd4582858ceccc5a8398ad0639ee7ccdb8cc46924c890007543b535066301ca

Observation 31f00f59-19a3-4c2b-a2f4-8491679d2c76 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.361604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.361604Z digest=sha256:ad4fe17365d505f972d10319da368ffa82995b5a423975c7e6cf378f9832a2bc

Observation f84149dd-0a3d-42b8-8d5f-c638c3e517b2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.530685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.530685Z digest=sha256:f71c73a38ce28238b6fe473b9a145c7533999f3acaee6be3c629f871f82d5de5

Observation 32fd6239-8a15-4a5e-8b43-1d4eb839d10e · outbound

This paper cites Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.676167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.676167Z digest=sha256:acca2876a4fa5ffb27437f5bfe9802f07377368d6c7561414d337f4d47749b31

Observation 13631173-84fa-4065-84da-7aa8bb414245 · outbound

This paper cites Private reinforcement learning with pac and regret guarantees.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Private reinforcement learning with pac and regret guarantees

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.913242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.801236Z digest=sha256:6c2a7b616de02d53b957e00deb1d53949ae09399bd068312463bc83e240a1fd1

Observation eb59928c-97bb-4b8e-80ad-1443f37636a1 · outbound

This paper cites TRL : T ransformer R einforcement L earning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment TRL : T ransformer R einforcement L earning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.749209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:01.936084Z digest=sha256:f5751c184085437d107554857024367f096e6804c22f83eadcd9558068566853

Observation 19ee0da3-f26e-4d8f-bbe5-0387e8bcc2fb · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.062867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.062867Z digest=sha256:73321c71214fd75693d45a2bbc78afbee3b5508e1e7dd8d38f4ba28b0126dfdb

Observation 6dcc33b2-c25b-4b26-9f24-290ac9944091 · outbound

This paper cites The Central Role of the Loss Function in Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Central Role of the Loss Function in Reinforcement Learning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.217073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.217073Z digest=sha256:09d591195b250be20f33dcaa96a99cb0e82988428a0a7fb5b2e963f1e38210f6

Observation 408a1957-edf2-4e6e-a5c0-60f311fc7d14 · outbound

This paper cites Oracle-efficient pessimism: Offline policy optimization in contextual bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Oracle-efficient pessimism: Offline policy optimization in contextual bandits

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.514295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:02.369874Z digest=sha256:521a7aea6ba540e0a88c3386dc079ff9ef38c8ea7f0fc6a305a85cda8d95a5d6

Observation bc6e284f-0f30-4db7-b4a7-e9324e6af1ae · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is RLHF More Difficult than Standard RL?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.561160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.561160Z digest=sha256:7287729b85c9d81f99b4ea435d0bf719ce759efabd81efa6cd0c181ef8dbf3d7

Observation ef0913f4-fe91-46fc-8f3c-89eb5bc4d9c7 · outbound

This paper cites an unresolved cited work.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.755041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.755041Z digest=sha256:02044f400457b8f97d7f94b76d208e599bd616e95428e93a25626bbb177e62cc

Observation 38d55713-5938-490b-a9f4-e55aa35bebe4 · outbound

This paper cites On private and robust bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.353478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:02.884175Z digest=sha256:220ef926db67de70f82a7a5b46d1e1d90da5f0718ab6852873b3f6895500b2d1

Observation d4a93fb1-1819-4d37-9e85-7a6155634ec4 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Self-Play Preference Optimization for Language Model Alignment

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.046755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.046755Z digest=sha256:13f35c55fc22658d65dced869b19cd8781a1019fe38de4e060a223a779afddc6

Observation 24c71555-cb6c-429e-81e9-b096df6f3489 · outbound

This paper cites On private and robust bandits.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On private and robust bandits

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:09.153167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.207476Z digest=sha256:d80f9d1fbf404f112c8cb8450d7e2c03bebc013aeff084e3326fec0a01824304

Observation 2d05d8e8-2c3e-4b5a-81ae-0e73ee18c877 · outbound

This paper cites On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.350482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.350482Z digest=sha256:04b53cfa29b1c030cf623963e3c889fa3052c4a3fa850829725b21e7f29b37e4

Observation c70c23d6-df86-41e0-8da4-2b2b588b97e4 · outbound

This paper cites Foundations of Large Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Foundations of Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.495496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.495496Z digest=sha256:9a9bd420ee832e2a8c5843f30e7dfa4db9ba6f14e6dd5abf243df51e877129ef

Observation 5f943956-9991-46a4-94aa-57ee8a6fb445 · outbound

This paper cites Bellman-consistent pessimism for offline reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Bellman-consistent pessimism for offline reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.964046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.635728Z digest=sha256:527ee3448581862d851b2a1addd888f3ed2c018d658936a1dfb344ba25cdeb1e

Observation 47693abf-7fe6-4f98-9725-ed925a50e51a · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.783637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:03.708960Z digest=sha256:12b9feb270d1d79499688bf6c1180b6d09db8930d4b490ea2a2b0e05ca90f06e

Observation 2c102604-556e-4131-b27d-df8e3dfe5770 · outbound

This paper cites The Role of Coverage in Online Reinforcement Learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment The Role of Coverage in Online Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.851174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.851174Z digest=sha256:91fbc9eb15e476ad8280a040dda26910e16ced93fb6c5ae73df10224385ba976

Observation f24e11ad-5bc0-4e7a-a0df-1b77114804f8 · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:03.971759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:03.971759Z digest=sha256:a56903a2f4d65c5115bb232b27714cefaea5fb71b1c1a8c561d48c6eb9d2bf2f

Observation 1e623a67-03b9-4747-8d5a-b1b369fbcd82 · outbound

This paper cites Differentially Private Fine-tuning of Language Models.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially Private Fine-tuning of Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.056564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.056564Z digest=sha256:0dec26672ba7e6eb1bd112e8f7c488dd3ae7eaf502b66cc246108c1a316128ee

Observation 99c7d7ac-d4bb-42ab-90b3-8e8cca818102 · outbound

This paper cites Offline reinforcement learning with realizability and single-policy concentrability.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Offline reinforcement learning with realizability and single-policy concentrability

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.527175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.163239Z digest=sha256:f8440dbca4ee0f633fb0714bdd1d92902b7d893f37edb52dd7d14849096aeaee

Observation 2a0bc272-c655-4f37-a4d1-9a8ef6274c35 · outbound

This paper cites D., and Sun, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment D., and Sun, W

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.315920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.365119Z digest=sha256:9b01646ff6ccb2faa555f61a22985cb654c6c8209d8bbe9bd743eb4d828cb54d

Observation 13b583d9-506f-4e74-bd85-4e04f870217f · outbound

This paper cites Corruption-robust offline reinforcement learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Corruption-robust offline reinforcement learning

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:08.096754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.456100Z digest=sha256:87e25053ff3e068d438daf1e6f7a152bcaf9a6d3c8fa6e633e224ede2cd33d11

Observation 0f25a33c-c908-4cf8-a2fd-d6b57a926c3f · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.554366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.554366Z digest=sha256:97caa8f62f744750211014c3d7ec217d2f3741e1b5c45f79f7a25a36cd7305fe

Observation b58fc734-8508-4caa-bbe5-520dddb3eb8d · outbound

This paper cites Locally differentially private (contextual) bandits learning.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Locally differentially private (contextual) bandits learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.834248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.710387Z digest=sha256:72b59259c72f71bb3e0f04f83063593e410328b97171a11fdf1a8067e4e275e8

Observation 4b5a4910-4606-41c6-a9f4-f1cab496e3d7 · outbound

This paper cites Differentially private reinforcement learning with linear function approximation.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Differentially private reinforcement learning with linear function approximation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.623841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.842225Z digest=sha256:23d44e298e4842f7bb0a365c709b46d6ebec420dce0f1e8758180a469a076342

Observation 8708e2fc-7628-4e8b-ad26-4173b24d4d05 · outbound

This paper cites and Tan, J.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Tan, J

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.448453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:04.955804Z digest=sha256:255b40550c8560a116cc258523e617f416ad10d7e4d27f179f43c31b6d083765

Observation a6de8d3e-efe2-4eab-93c7-daf7f2c38802 · outbound

This paper cites and Zhang, W.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment and Zhang, W

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:45:07.202397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:45:05.082364Z digest=sha256:f5319868d612d20c98c9c76352c2eba33ea688aa619a81a074e9c9d142fb3999

Pith citing papers

No inbound Pith citation observations are available.