Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

As of 4 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2606.00367.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.00367 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:13:07.480013Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5cb98f12-076b-4673-9e14-f2fa72c0f437 · outbound

This paper cites Shah and Martin J.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Shah and Martin J

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:5ff93fe7b2fe09acaf95ea276364a4bdae2304c357e6ae05ea7e053388b2cf12

Observation e38bb2ac-ca24-4309-981d-2bb97ce077af · outbound

This paper cites Journal of Machine Learning Research , volume=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Journal of Machine Learning Research , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:44eb97d012afb56afee63a1e8e6263a1ed241a7642e90d5a35e2e822f743ea54

Observation fe76f9a0-c91d-45c9-92f4-9cc96da2f170 · outbound

This paper cites arXiv , year=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems arXiv , year=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:8f42a4a0b1d504df43be4868b1d6fa9302b10ae253a02837ab5862da3879219a

Observation 26cd1e61-eae6-4725-bfc0-5ae2f938bab0 · outbound

This paper cites May , title =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems May , title =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:e2d708bc19252d380afb753b9704c4270487c7ca0de4116253e50ee0def5c609

Observation b5663762-6719-4f93-aec8-51e38c2d6fc0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:3174309ff88583034e069718b41d77f4861c289de23b52a11a87b75ceb68f0f3

Observation 6041a12a-e0a6-4968-94b4-712a68ce32f4 · outbound

This paper cites Theory and decision , volume=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Theory and decision , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:8b31b5f64465fd82910583d050044af3091984de6db90c4564a56a0d6ccc24d3

Observation 58c5b4ee-1002-432e-816e-b7531e30475e · outbound

This paper cites Transactions on Machine Learning Research , year=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Transactions on Machine Learning Research , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:736e8fda5d8800d75b31696cd20141ae73c92c8769534cd3a98679b937f3683f

Observation c5bedb28-f743-453a-b0cb-49f71478f336 · outbound

This paper cites Reinforcement learning from human preferences within the.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Reinforcement learning from human preferences within the

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:f0de99f8ff5aed644d95ebe4bac6b186952f435469bb76e01c15d98c5806a029

Observation a78ec99c-cb77-42a1-a4e8-e093551be050 · outbound

This paper cites Papadimitriou and John N.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Papadimitriou and John N

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:b8052a0d4ef7f7700839f8e7f1dbe870e4b0028cd2f40cbd925f8e0b3f0b2bba

Observation e42fce6b-881d-4846-9a45-12dd4e26bb80 · outbound

This paper cites Advances in neural information processing systems , volume=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Advances in neural information processing systems , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:001fcad9c719724049fb05986f8b0fb4d67c1f31363b934b11af991802316400

Observation 369e9412-b1e9-4475-8a90-2495e08c878a · outbound

This paper cites and Van Roy, Benjamin , title =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems and Van Roy, Benjamin , title =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:af2782d54302fc911f37533e66382f08687f1fab5f6582b3003357bc006633de

Observation 89237b8f-e6a6-44eb-a3b5-947694057ff9 · outbound

This paper cites A decision-theoretic generalization of on-line learning and an application to boosting.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems A decision-theoretic generalization of on-line learning and an application to boosting

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:d62ae0c7c02b9a52568da433a15ffd44ae78521c5b9197eadcacad67e64f24f9

Observation 8d3117e8-9a6f-4f0e-9be4-42450d8d7d7a · outbound

This paper cites Transactions on Machine Learning Research , year=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Transactions on Machine Learning Research , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:1ef32cd9b02b402ab3de8c8f6cdf2687e4a05ec32023e32e9f199ed38e0bc500

Observation bada31c1-e44f-40ec-ab28-fdab3b64f101 · outbound

This paper cites and Lowe, Ryan and Voss, Chelsea and Radford, Alec and Amodei, Dario and Christiano, Paul , title =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems and Lowe, Ryan and Voss, Chelsea and Radford, Alec and Amodei, Dario and Christiano, Paul , title =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:a752f19f822951ddf046e2db654db2e641ebcd5bab35f7cf5eeeea3189d92ad7

Observation 0c39a32a-d782-44ca-b8e0-b2c7355551ca · outbound

This paper cites International conference on machine learning , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems International conference on machine learning , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:771f8d4cddb64484991d8ee0c979b7ca8001e543e498fc28b7c1b74accef6965

Observation 72897258-17e2-4c1d-9206-d7a7a02ae573 · outbound

This paper cites , title =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems , title =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:472318096b4b87857a8bcb9ddb84ecd23a24570fe9be526ba2a498966180d4b2

Observation 15bf99a1-2a47-49e9-99e7-5d7ed4ceaeda · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:1570e72ea3883652fa3944e51797eb72cc0ed8fef2575d613534a8a57c5fd9b0

Observation 6840f376-735a-403d-8706-9cd30ea61376 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:2251f9dfb8d0bd40be64fb2f8bbc21cd6ccebc79534f12459a7da3bd4d4a3ec9

Observation e937cf59-00c9-4f58-915f-ba8f0c385e3f · outbound

This paper cites 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:4d8c72de3e0ed27212a366e15bd48efe90a8fb4cceeb2334e6b448f1542d3d5c

Observation 385ddec5-7ef5-422c-8334-e7b13d79a1a5 · outbound

This paper cites 2025 , booktitle =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2025 , booktitle =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:ab2300d590e66966d49a7347dcc3e78f01635e18078a76355332e20c1c9bf538

Observation f8bedf18-9abe-4c4c-b2ee-2986d93bc196 · outbound

This paper cites Online Markov Decision Processes under Bandit Feedback , volume =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Online Markov Decision Processes under Bandit Feedback , volume =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:aa82d645857383aafd09ef9525b208f5b354bdd23d4686f5d08047049720ccb2

Observation b8009213-e42e-42eb-9d71-a4054019fe36 · outbound

This paper cites 2025 , eid=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2025 , eid=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:3b605ada9a4dc47ed0851c059f367cdbc832a7ad8fc901afe2999830afd36917

Observation f0df10be-09f6-41ea-8a80-8facc9461300 · outbound

This paper cites Econometrica , volume =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Econometrica , volume =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:3564ace17afb2b14150a19771e18950c0f21d42d1b0cf2c0d815d0a508190efd

Observation 407f1cd8-225d-4db0-b97f-3384d5832503 · outbound

This paper cites 1982 , author =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 1982 , author =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:07ac45812b50b9a0b354ef28fb0eab90a2fd75c404a4974ee32209919ca65226

Observation a6e77f40-0154-4ab5-a1cc-e22a27d8e775 · outbound

This paper cites 1984 , author =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 1984 , author =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:5b61ecae5fb38b0e4bd7495db424d6226367b29444f939daecb8c1e2c2b70eb3

Observation 0bfbc581-7c10-4682-b07f-f0c49698fe76 · outbound

This paper cites Journal of Machine Learning Research , volume=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Journal of Machine Learning Research , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:2c0d691f8111672464baaabf4db600a93909d6715cb9b2d2338eb556e896367c

Observation 239aae29-1fb5-4f37-ad83-9a2c664f4858 · outbound

This paper cites 2024 , url =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2024 , url =

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:ccafe2af95d0fafefc05493f5a0a09a26668e0996506004c666a22226770dd18

Observation f305c3d4-ab44-42cc-adb3-f58472c0cb94 · outbound

This paper cites Prediction, Learning, and Games , publisher=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Prediction, Learning, and Games , publisher=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:005f18f4fa1aee5fa0ea4dfc9163ce1c0c5c82a11979ddd1dfc75aadafcebc85

Observation 89bb9e37-08f8-40af-84f8-320f53b17513 · outbound

This paper cites Pacific Journal of Mathematics , volume =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Pacific Journal of Mathematics , volume =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:dbbeec148e4e4401b25fcc6c93b5622d2a8e7d069a71d7a00c11a19b742b7857

Observation 1355fa2e-9810-492c-9a2e-675667efd0ae · outbound

This paper cites Theory of Computing , volume =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Theory of Computing , volume =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:25f8bc1952308505803aef1ed462b9eb0600b0c00e6a8b0dc109068ee9798207

Observation 176c218c-c04c-4d93-b19d-70d86b0f4bf3 · outbound

This paper cites Araújo , title =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Araújo , title =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:bdfe6dcdfafc3816792a8524ea408b1f47eafe5243857223a784bf5037fd8e5e

Observation bc2f9af5-260d-47a1-9ed8-2915d8dc6f08 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences , volume =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Deep Reinforcement Learning from Human Preferences , volume =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:2d784320f7dbb65af81a7d2b84e960761e91e8c4ace8c00e93baeb44770e61c7

Observation 67793272-22ff-4676-9f88-c0779948146f · outbound

This paper cites , journal =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems , journal =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:1069be53d0756c180b63c5ff8caebce8a225d72be2591177f5b7ec0063b398ac

Observation 2ce2ddfb-1c1e-4f3a-85a6-4be06e2199ed · outbound

This paper cites ICLR Blog Track , year =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems ICLR Blog Track , year =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:3daadd7fe44fc634b9d0d8550df6525fa090b2501dcca6da1d82c92e17c9e3c9

Observation 8f33edf5-ea64-443b-a769-003ae422136e · outbound

This paper cites International Conference on Learning Representations , year=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems International Conference on Learning Representations , year=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:6b0d4c0884e54d5cde92042fd8702f6e35b2691d9ed31d57e77a4c6087f211f7

Observation df928353-6c78-4903-a45c-1dd1155b65b7 · outbound

This paper cites Iterative nash policy optimization: Aligning.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Iterative nash policy optimization: Aligning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:95c9ed36b2c511bddd3d7dc8966d4ec6d57b9de727ec7f1ffae12e4196a18036

Observation 3f14cae9-e292-4ac1-912f-61f91fd583fb · outbound

This paper cites International Conference on Learning Representations , year=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems International Conference on Learning Representations , year=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:2229ef4ac733b7c6655f0695f7bf6fe316cdbbb6118708227ec010e779ef179d

Observation f88d88e9-0f91-4b65-8b9b-0143a0fea496 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Advances in Neural Information Processing Systems , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:55540839d73aeb9a50ae0dd20c42776e22ce516a4387b7152bb78440c538cca3

Observation 1022f6f2-b343-4f32-ae25-905f92713a3c · outbound

This paper cites International Conference on Learning Representations , year=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems International Conference on Learning Representations , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:447f2b570e4fbb37ceb6b7a058c5c1905c0a46c9dcb179f723e9cd5c83a196ea

Observation b6aefaed-b482-47f2-ac07-3ee7c2faa879 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:c56848d38c15372a9607d7b139532a7c243c10155715f652b525ade058b03a69

Observation af831697-eba1-422e-b551-afa610766027 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:0d52e4fb0a38906d7019ad7e3d85f01bdcb018e8cc01c5fa1d2e2cf1110180f8

Observation 2e96537d-b3ad-4b77-b04c-0cecfb2a6152 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Proceedings of the 41st International Conference on Machine Learning , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:04ac311db1c2e2c68e48dc7697592b2cb13f05f577015e59fd257bda08939a86

Observation 04a5b9d9-e1d7-4c48-9015-32b4cee2f23d · outbound

This paper cites 2015 , booktitle=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2015 , booktitle=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:1fdeae21a1d70257d26bc49e0786e87d36e8b3c3cefae3d22801268d17dad78f

Observation f3a0b3ea-f36f-4139-825c-3676fe4cc2f8 · outbound

This paper cites 2016 , booktitle=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2016 , booktitle=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:b23d06c07ac1e25a139512a1fe36f1373451b893cdb89bb26cb4e1b8ed6b3aa0

Observation b2bd2f4d-5513-4012-b1af-3afdb5e9d07f · outbound

This paper cites arXiv , eid=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems arXiv , eid=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:8399553b645254805f1f240dfd133ff3ea83a4e53d0e3be406aa57483f6fee24

Observation 0a116ac5-3549-410f-b008-d0c919bf1ba7 · outbound

This paper cites Proceedings of The 28th Conference on Learning Theory , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Proceedings of The 28th Conference on Learning Theory , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:35b1ac926997f1b26895e27036a64d4cd390ec0ad2828e1703c4d47d5a0b146e

Observation 136f7d0a-4797-4591-9fd8-ec4884eac504 · outbound

This paper cites Proceedings of the Twenty-seventh International Conference on Artificial Intelligence and Statistics , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Proceedings of the Twenty-seventh International Conference on Artificial Intelligence and Statistics , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:e4f4f612e8f5b2ce09c4ec1e42a85c8471567aa00345adb45223c477d68ae9df

Observation f2453679-a98a-40f0-9e0c-d84a3ce0197a · outbound

This paper cites 2014 , publisher=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2014 , publisher=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:bcfd06b0acb4e2e3fb2e359829894890d8252a3ae2d83154a6bfc4329a775059

Observation 3142b86a-bb7b-47ea-96c0-fff7883266d7 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:2a0d5a08179e909a8b85947a6d80a7ce3e00dcace2735f53dabea46337ed7aba

Observation 785776f9-5346-4058-a937-e8fe2384e564 · outbound

This paper cites arXiv , eid=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems arXiv , eid=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:c4e36591a295555f3119d124fb192ff55d3b19b71cae389d76d94a92f4e0c498

Observation 02d45875-3276-49f0-aaff-a1db0c12a184 · outbound

This paper cites 1997 , author =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 1997 , author =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:919b5406ec461bb24ba7e00a3e042ccbbb2645d90e1dd72ae611893b150879e5

Observation 8c292762-782e-42a8-bcde-9a3a2b1029c8 · outbound

This paper cites 2008 , author =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2008 , author =

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:ce093b3dadec7658a849f6107c1fa570abca2005dec0dde890824ec85585bd81

Observation e77d0791-d7d3-4d66-88d5-4b3deff0bac1 · outbound

This paper cites 1999 , author =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 1999 , author =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:4cf197aff273413c67ce3f54c13f3de39c32ceb8edabe7a892c82071ef55b0d4

Observation 8b0727b4-0719-4f6b-ba95-ffd1a16a01ea · outbound

This paper cites Proceedings of the 31st International Conference on Machine Learning , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Proceedings of the 31st International Conference on Machine Learning , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:1eb2e64bdd304f4564cae9307e5b848d50059ccb13941458561962702dbfd745

Observation 8dd627c9-62ae-4831-ae2b-0466744ec6a3 · outbound

This paper cites International Conference on Learning Representations , year=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems International Conference on Learning Representations , year=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:bfbf0191d3ecbfcca6de37dc5158b4554adfa4bd3a0e10bca724e0090e5c150b

Observation f5667c41-439e-4f37-be55-b497ca8a0267 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Advances in Neural Information Processing Systems , volume=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:5be8013a107a98a7ba915f2c3aa24dfdaa093d546394e36c1de5b77f3bbdba2e

Observation 16e3ce12-fe2b-49c8-96cb-f5dfa90f0611 · outbound

This paper cites Proceedings of The Twenty-sixth International Conference on Artificial Intelligence and Statistics , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Proceedings of The Twenty-sixth International Conference on Artificial Intelligence and Statistics , pages=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:be02358c8e07a85486ceb68164e55b34b3188a1122eb93a57afce7b155df7d71

Observation c145db34-3c16-44b5-aebc-ba5729c9e32b · outbound

This paper cites 2024 , journal=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems 2024 , journal=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:1b2704f28abe151022166b2df5d818c33e296879bc01cd1b9ed4d884b4d4b9a6

Observation e14bfaaa-6504-4c1b-a462-84ade81ad2cc · outbound

This paper cites International Conference on Machine Learning , pages=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems International Conference on Machine Learning , pages=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:2d67ed9820240f793a8e6c8026b882b136fee9eecb85578ac90c43fcdfed8a82

Observation 0536c427-bf36-4051-9d0f-4ed1d471318d · outbound

This paper cites Jordan and Joseph E.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Jordan and Joseph E

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:d3e0df1d6cedf7bae4374b167f835b31a11ad908475f62a1bee20317de6d4e91

Observation a3cca256-fe0f-43ff-92ec-78ef143bc704 · outbound

This paper cites and Barto, Andrew G.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems and Barto, Andrew G

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:276417e09c7990892f4a90b4aec7dbac06b3ac431921f1bd51022b7c07fd225e

Observation 2619c0c2-f68c-4c18-93b1-d4406cc3650b · outbound

This paper cites Proceedings of the 32nd International Conference on Machine Learning , pages =.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems Proceedings of the 32nd International Conference on Machine Learning , pages =

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:6076a1e7841c3b451b32cc4090067f64c7db7924f72ee7365efac37ee886e4c9

Observation 9b18b520-aeae-4bd4-a653-8c884729df8b · outbound

This paper cites arXiv , eid=.

Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems arXiv , eid=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-28T23:13:07.480013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T23:13:07.480013Z digest=sha256:b6db4618db11d45d9144c1399a29c5ba7df271dbd9f32990463ac3001add3954

Pith citing papers

No inbound Pith citation observations are available.