Pith. sign in

Paper Citation Record · LEDGER

Combinatorial Reinforcement Learning with Preference Feedback

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 1 inbound Pith citation observation for arXiv:2502.10158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10158 v3

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.158975Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:57:04.381612Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T19:57:04.556666Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact2
  • verified fuzzy42
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b18648c4-8d23-482a-b0d4-776f4aea6e69 · outbound

This paper cites Instance-wise minimax-optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Instance-wise minimax-optimal algorithms for logistic bandits

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.791208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.224384Z digest=sha256:0f93e6738c16cb919250979f9bf95359a284d5a95528d2d6e7df258562c229ef

Observation 28165428-c5d9-4984-b1fe-ef7bca76f6de · outbound

This paper cites Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.761100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.231128Z digest=sha256:d64d51e715a44dacf7afaa049defdc86cdbbe769e915078979f6e372529c9377

Observation 2f556e53-d9b4-478a-93e5-6bad9c169184 · outbound

This paper cites A tractable online learning algorithm for the multinomial logit contextual bandit.

Combinatorial Reinforcement Learning with Preference Feedback A tractable online learning algorithm for the multinomial logit contextual bandit

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.731998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.237577Z digest=sha256:504e40cdbee3a1bdc79a4c6514fa774a54bca180bafffa961c646256c7f86596

Observation 548ffc32-b4db-40e6-afa2-84473a54d7ee · outbound

This paper cites and Goyal, N.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, N

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.706758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.245239Z digest=sha256:e420e65c2ceff3e38ca6026d6589c51b0841e6db6719a0b0f3a55f08551ecaa4

Observation c5559935-c99c-4704-94fc-f6e73a4e3652 · outbound

This paper cites Thompson sampling for the mnl-bandit.

Combinatorial Reinforcement Learning with Preference Feedback Thompson sampling for the mnl-bandit

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.681536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.256779Z digest=sha256:b3440a755873664e1867f556bb7f21ae5405f419478aeb13c1aa244ecee80f73

Observation 0080f2ad-5033-432b-92f4-b3169a7009ab · outbound

This paper cites Mnl-bandit: A dynamic learning approach to assortment selection.

Combinatorial Reinforcement Learning with Preference Feedback Mnl-bandit: A dynamic learning approach to assortment selection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.657502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.262831Z digest=sha256:eedef0bfcbc280ea71a13d5ba08d546bd9f1aacb01ba9e753c4759899ee579a5

Observation 1a6b6990-e2f1-4953-b1bb-3f9504f68356 · outbound

This paper cites April: Active preference learning-based reinforcement learning.

Combinatorial Reinforcement Learning with Preference Feedback April: Active preference learning-based reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.271517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.271517Z digest=sha256:f5e7a98d2abbf2da84e02a2ee8ee5884e2f67d8472a575bcd24078855e076668

Observation 6ce715ec-61b0-468f-9f9a-9d88960490eb · outbound

This paper cites and Thrampoulidis, C.

Combinatorial Reinforcement Learning with Preference Feedback and Thrampoulidis, C

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.611297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.284760Z digest=sha256:38e4dbb5be13375c00a39d9fc6c7a6b3e1c80b5016572248d4724c5d283daaf5

Observation 87e0a6b4-94a7-4b9e-af07-8a22e8837bc7 · outbound

This paper cites Distributional off-policy evaluation for slate recommendations.

Combinatorial Reinforcement Learning with Preference Feedback Distributional off-policy evaluation for slate recommendations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.587861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.292863Z digest=sha256:937b5798df6b96ae03d03effaeed9af78be2595c7cc50c7cd08d12946acadb40

Observation 280622c9-3417-4960-854e-bad58b1406bd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.561019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.304989Z digest=sha256:65dab2dd9e1c82d8bcfdfb00449e1e5600def180bbd2753e2e5e5a97efdf9cb8

Observation 2ba47c7e-b279-4d54-8431-699597d2ea37 · outbound

This paper cites Combinatorial multi-armed bandit: General framework and applications.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial multi-armed bandit: General framework and applications

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.532539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.312614Z digest=sha256:29c21efbb41379f75c972993e8ec026010948299d4f55c8ecd1e3a69c9757e9b

Observation f373dd7e-7ebb-408d-b664-499d8c7959c3 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.507919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.323399Z digest=sha256:e4834a0964d11281a7a0e8fb7730a1fe4000e6f1844b2ec219d28ff20542f1bd

Observation 1301a24b-9616-49e4-be44-00e3b0eebef5 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Combinatorial Reinforcement Learning with Preference Feedback F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.332237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.332237Z digest=sha256:4e325aebde926c01c608ef80410fdd23bd6c1e216519915c4b630dedbe4dc61c

Observation 9f827c34-880e-4eb0-b828-1d6e72e67e0d · outbound

This paper cites S., Proutiere, A., et al.

Combinatorial Reinforcement Learning with Preference Feedback S., Proutiere, A., et al

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.443014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.341524Z digest=sha256:888c321bd5022a9b4df75c8fbbc3535c78e75f089787d177f94b59810c1e45d5

Observation f486ebc0-5827-49a4-b03e-e77ade3febf1 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.415420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.352173Z digest=sha256:6d7f28dd7331d2d15ecbbaff5369b72c9a78b24e6cca16c4dfd720cbcaa442a5

Observation 6efd8f5d-8ecd-4f79-a205-07b39d9a6b00 · outbound

This paper cites Assortment planning under the multinomial logit model with totally unimodular constraint structures.

Combinatorial Reinforcement Learning with Preference Feedback Assortment planning under the multinomial logit model with totally unimodular constraint structures

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.383925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.359505Z digest=sha256:c9cbd3ac5bec02a23e21737486fbab74d0d86ec24a450715ec0015416f41de9e

Observation d28026fd-a4ba-4d6e-99d7-21851d4eb11f · outbound

This paper cites Reinforcement learning with combinatorial actions: An application to vehicle routing.

Combinatorial Reinforcement Learning with Preference Feedback Reinforcement learning with combinatorial actions: An application to vehicle routing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.337692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.365124Z digest=sha256:f074b44a93077c105c62fccfd180417b5a7424663f2362e85641989d39f7968d

Observation dfac6861-d89d-44c1-9e18-c0cfc3117ede · outbound

This paper cites Bilinear classes: A structural framework for provable generalization in rl.

Combinatorial Reinforcement Learning with Preference Feedback Bilinear classes: A structural framework for provable generalization in rl

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.314252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.371484Z digest=sha256:98eae26e0af12ad6d85a0f660fdc11aa67c779a31f5171a736bb0c64e47362ad

Observation ac33eec9-1542-41d9-aa9a-81687450b21a · outbound

This paper cites Cascading Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Cascading Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.850736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.377776Z digest=sha256:d9d538b864c49146c571e6c7a1364d640caf79fb6f1483f2fa60595e7f7b2a75

Observation cfa91a6f-ffe0-4abc-946c-00885a286f88 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Improved optimistic algorithms for logistic bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.405443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.405443Z digest=sha256:2942560207a2ad393b8d0ed48ee785dd2d8c688fc9545f04019a8b391d6867db

Observation 098c220a-6abd-4050-b907-6c19d2a35cb0 · outbound

This paper cites Jointly efficient and optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Jointly efficient and optimal algorithms for logistic bandits

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.269061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.475614Z digest=sha256:09d94cbf3e2d0a1dac48de79e781b9e0c163e9d548dc25f0de1cecc49f2d56f6

Observation 7e340fbd-91da-4e35-ba4b-b239c9866093 · outbound

This paper cites Parametric bandits: The generalized linear case.

Combinatorial Reinforcement Learning with Preference Feedback Parametric bandits: The generalized linear case

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.243063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.546882Z digest=sha256:bf1db5f69e80a51913c52825341d68ea45f2fe62a2623e990994ba3f8f973227

Observation 2f8a65c8-fb8b-4f3f-8849-0f9032f82e28 · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Combinatorial Reinforcement Learning with Preference Feedback The Statistical Complexity of Interactive Decision Making

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.598167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.598167Z digest=sha256:c4d870dd08e9bbf46054f02866b6859c83370ae12a8f3bd6a9a24187c11a3235

Observation c03b3cd5-d973-41bd-a51c-8dfa139917cd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.218068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.655974Z digest=sha256:a19e99e37470e9a7cb1ee17d88b43f16d06282e4253c7ce41e90fa77aabb2384

Observation 6928cb0d-781e-413a-8e6e-441d96001d7d · outbound

This paper cites Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.748595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.665758Z digest=sha256:03c1f99cfe56b916f62088516606ea7024724964eb3fc6124cbfe3de79a702a9

Observation 3d67be2f-445b-4ee9-8b72-fb62b692493b · outbound

This paper cites Slateq: A tractable decomposition for reinforcement learning with recommendation sets.

Combinatorial Reinforcement Learning with Preference Feedback Slateq: A tractable decomposition for reinforcement learning with recommendation sets

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.197624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.672769Z digest=sha256:e943d414abf4332fca206b81b28c9d8f966bdb8dc319a8a41d6556f9262e8e24

Observation db054e73-4680-4d27-87cf-20070c8075d5 · outbound

This paper cites Randomized exploration in reinforcement learning with general value function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Randomized exploration in reinforcement learning with general value function approximation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.680375Z digest=sha256:d7cd30b6ba96c5fa9a96e4f5e17922579e58abbb19f34fcf0cb84ec66113cbd6

Observation d37ac1b4-fda5-476f-a172-b224f7503961 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.133885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.685870Z digest=sha256:eb0ffe296af3961fd7c8784e856efa4d9eb242488f5ac1aee02d35261fe0f839

Observation eb345469-8408-4da7-ab13-a4e73fec3f43 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.102859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.692642Z digest=sha256:95e85b058f8f5971177b831dc22df485801325bdf7252db9fc296c335fbcd100

Observation 768a359f-e859-454d-8b5d-64a5cc65c0f9 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.076408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.699888Z digest=sha256:adeca0c34a9d6ff9ccae9269572c746b9d086123cbdc30c3166099d5743e6848

Observation 20505fb5-e7b5-4bb6-bba8-f729ebe59dd9 · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Combinatorial Reinforcement Learning with Preference Feedback Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.053824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.705775Z digest=sha256:e0ec5cdca619b548577db1cc98eaf0d83a7bd9c5f9d401e6a535a7b6e3c95f5b

Observation c7559e90-d365-4d5d-887e-5c33a19d1edf · outbound

This paper cites Online Sub-Sampling for Reinforcement Learning with General Function Approximation.

Combinatorial Reinforcement Learning with Preference Feedback Online Sub-Sampling for Reinforcement Learning with General Function Approximation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.711779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.711779Z digest=sha256:f495207e6bee756624a648f8a178e78d7566db82d1fe7c484ed9f7fd275357e1

Observation e51f6d31-7122-487c-b078-3f5864f46902 · outbound

This paper cites Cascading bandits: Learning to rank in the cascade model.

Combinatorial Reinforcement Learning with Preference Feedback Cascading bandits: Learning to rank in the cascade model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.022919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.717594Z digest=sha256:f3166ca32004e48af382009bde78a9c92e14b2ac260767a844542bb4aa31c563

Observation 97c5c92f-6b79-4add-a98d-4cb47ffe6888 · outbound

This paper cites Combinatorial cascading bandits.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial cascading bandits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.990933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.722642Z digest=sha256:099212f8a4cea5551f5cf86c978b46ea4f1160412084f14e39ee154e1a7122e0

Observation a616e48a-f068-49d1-b623-93d9fbe07a2e · outbound

This paper cites and Hutter, M.

Combinatorial Reinforcement Learning with Preference Feedback and Hutter, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.969336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.727018Z digest=sha256:fffe97f6b784a1343baa914824e2630971e459530c9231bbbb86f428afe6d325

Observation 1d65c601-5eb0-47a9-a4f8-b27ae658b390 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.946498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.732283Z digest=sha256:19fb11dd5b5bd0b7eabee4b3fde7893c194ed6d68c93f5c76b7a5ebd7cf85c56

Observation 28b636af-c5f1-4f43-896a-d20548e80e13 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.921457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.736731Z digest=sha256:97b4fd6e17863a6e1231a8abe102560bd54614e43f549c866ba529f81ee0ed0a

Observation aabc5773-222e-482b-97a6-cc8434b97888 · outbound

This paper cites Online learning to rank with features.

Combinatorial Reinforcement Learning with Preference Feedback Online learning to rank with features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.741426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.741426Z digest=sha256:619e6042735db1f7e20fc0af3f8d2c7a12fb2e53c9be229d5941ec2eaa4c7b65

Observation 2e7da764-f091-4cd8-8af8-a1b6b9e12d03 · outbound

This paper cites Modelling the choice of residential location.

Combinatorial Reinforcement Learning with Preference Feedback Modelling the choice of residential location

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.869998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.747449Z digest=sha256:095cb58fa172cb835072d481c8bac614616794d3277bc202eef0a3d2717276ac

Observation 7ea4aeb3-84a7-439d-8dd2-8f0fb9d3985f · outbound

This paper cites Counterfactual evaluation of slate recommendations with sequential reward interactions.

Combinatorial Reinforcement Learning with Preference Feedback Counterfactual evaluation of slate recommendations with sequential reward interactions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.846788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.754104Z digest=sha256:6115d3ebed7390f615b878baf82371c3b2908488d965db2d9c0e5aa51abd955b

Observation fcfff1cd-f4ab-4158-ae4a-b06c3f79319c · outbound

This paper cites Discrete Sequential Prediction of Continuous Actions for Deep RL.

Combinatorial Reinforcement Learning with Preference Feedback Discrete Sequential Prediction of Continuous Actions for Deep RL

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.761945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.761945Z digest=sha256:59440b6839bc0cc6c8939ab0f61a0e8014ae986dbb8dcbf1d2cbcebcf00d78fd

Observation f28ccee3-682d-4983-af4c-fc167647d0c5 · outbound

This paper cites M., and Van Erven, T.

Combinatorial Reinforcement Learning with Preference Feedback M., and Van Erven, T

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.823133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.792670Z digest=sha256:a1f78e599398a2d7fdc041a07ed7c7cb4533d98ace9a2bc6528a156cd91dec22

Observation 56c2cf88-add8-4503-afae-f42231ec6f72 · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.798459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.825505Z digest=sha256:8a8b8cb60bc8eec5aaa9eeacab2447ad0ab64b1acb8651b6ea72be9bce1e2f15

Observation 817c122c-11e5-4d5e-a8d0-f173f9c3e92f · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.776371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.844502Z digest=sha256:7b0d85991c90974b52a2fb2a2de8a88ddbae2bceca6eac78896070f14c6dfc1e

Observation 5709ed03-ef9c-4ebf-8579-29b622ec9063 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Combinatorial Reinforcement Learning with Preference Feedback Online Learning: A Modern Introduction Using Convex Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.870065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.870065Z digest=sha256:8762ad75717672c84cdaa7e21bf11aeee9ff4e6fb5db0edd83d298cfc16fb1ed

Observation 71557d63-5f85-40df-aeba-389f1c2d1a42 · outbound

This paper cites Training language models to follow instructions with human feedback.

Combinatorial Reinforcement Learning with Preference Feedback Training language models to follow instructions with human feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.896163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.896163Z digest=sha256:fab6a1d277c424b07bc48fda49c9e74a1f631acf3b5ee84cd889d46abda33f55

Observation b44bb0a4-8212-49ba-821f-f13f5e3bc851 · outbound

This paper cites and Goyal, V.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, V

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.725451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.904201Z digest=sha256:b1a9e42b8cef5bc237cdd783ad198f36f65ae7ecbf5be38dfc6cc90299764f19

Observation a975f62f-39bd-43ad-aa70-cfa62807d193 · outbound

This paper cites M., and Shmoys, D.

Combinatorial Reinforcement Learning with Preference Feedback M., and Shmoys, D

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.702565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.910945Z digest=sha256:aa19c254475734b9100837b2d52b3e3df846943b54a08f66856b01b212e7c7ab

Observation fd73fb0f-6618-47e3-8b24-2fa8b5fc3fcd · outbound

This paper cites and Van Roy, B.

Combinatorial Reinforcement Learning with Preference Feedback and Van Roy, B

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.674460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.918859Z digest=sha256:f84e8d4df6d44d8a3155fff2c4af76a15a9c2ad1ccae2c4e4b44467df65a3402

Observation 3b9cb725-dbc2-4f30-886e-b86fde07259b · outbound

This paper cites CAQL: Continuous Action Q-Learning.

Combinatorial Reinforcement Learning with Preference Feedback CAQL: Continuous Action Q-Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.926888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.926888Z digest=sha256:0765100321ed1703d95f522fd7e85001b458f0f1a54c524c2ab3c4635967df0b

Observation c361c8e7-8103-4598-9c9c-54a000858ecd · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Combinatorial Reinforcement Learning with Preference Feedback Dueling rl: Reinforcement learning with trajectory preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.934451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.934451Z digest=sha256:f74635bfb47cc86fe3b89422645035c67c1c5cc942e82077efb2dde03a9a2c88

Observation 0f14d212-1104-4eb3-a749-0100b3ac460c · outbound

This paper cites and Zeevi, A.

Combinatorial Reinforcement Learning with Preference Feedback and Zeevi, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.571717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.943188Z digest=sha256:42c853d1501b78577a1aa22a6ba2d9e514fbbe49b9344e92c64eaba8fc6f3dea

Observation e54c55dc-8a4a-4da0-b233-80cb91fa2068 · outbound

This paper cites Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.951927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.951927Z digest=sha256:60362a013ed5ebd667ebca8b25e2c9f805601e16f0781d3ba7c9718338b43983

Observation 2d6c0153-9fd0-48af-aeb0-753d3caa5d61 · outbound

This paper cites Off-policy evaluation for slate recommendation.

Combinatorial Reinforcement Learning with Preference Feedback Off-policy evaluation for slate recommendation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.490085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.958603Z digest=sha256:9fe189704c561a7c3ae9f6ac2bc729ba160e1ad309108c80859468b7513ac9b7

Observation 966e96de-4d75-4ed4-a76e-7b8f27233eb2 · outbound

This paper cites Composite convex minimization involving self-concordant-like cost functions.

Combinatorial Reinforcement Learning with Preference Feedback Composite convex minimization involving self-concordant-like cost functions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.392696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.967403Z digest=sha256:4b83b1a0f0a3218ff46ef4f2821a0d49888485eaf07759e29c9cdfde684ccefd

Observation 0381e4e0-bdaa-41eb-8fd2-05ade2ab0bc1 · outbound

This paper cites Control variates for slate off-policy evaluation.

Combinatorial Reinforcement Learning with Preference Feedback Control variates for slate off-policy evaluation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.370807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.973733Z digest=sha256:7d72f44472180732d0015563db3276d458fa0ba0d5064aa0a247029a8ee4ff55

Observation 94a32340-6a5b-4085-bb4e-f7ca283eb595 · outbound

This paper cites R., and Yang, L.

Combinatorial Reinforcement Learning with Preference Feedback R., and Yang, L

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.343071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.979355Z digest=sha256:8e6337fc2a6e8acb8edf5c04d65a4467ee8408d3c3f9b9067c83f3186a34a6d1

Observation cf765a08-75ae-4f55-8dba-ffcf813dc963 · outbound

This paper cites S., and Krishnamurthy, A.

Combinatorial Reinforcement Learning with Preference Feedback S., and Krishnamurthy, A

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.299861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.984567Z digest=sha256:a71c5cc3412cca67139644f5b8ecab05649d017fa8fb2d19fc1c049432920938

Observation f7cfa02b-79b3-4be3-9004-935ad45b2f17 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Combinatorial Reinforcement Learning with Preference Feedback A survey of preference-based reinforcement learning methods

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.992315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.992315Z digest=sha256:2532db285a4dd088521dd871d99ab59e75493446f3cc0f13bc8b647be6ff14c9

Observation 709c594c-281b-4b1f-b016-6bb1a69643b8 · outbound

This paper cites and Wang, M.

Combinatorial Reinforcement Learning with Preference Feedback and Wang, M

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.250244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.998266Z digest=sha256:b6e4e1b626f9c3a1282c14555a3d005b0f76e68bcbd538f67a38ac604dad927d

Observation ba8ba420-fd18-4f34-b324-9b965813f310 · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.002840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.002840Z digest=sha256:9bfdb4abd31abb13df45ac36b625322ac6bbf9654fc3ea5b553dc645b7c3f6dd

Observation 304bad2e-afe8-4a0c-86ac-823bd00b71a0 · outbound

This paper cites and Sugiyama, M.

Combinatorial Reinforcement Learning with Preference Feedback and Sugiyama, M

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.224992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.009631Z digest=sha256:3a0d6ed4ffb10be6fec9a25a2de0ee4addc6d71fbefecf5216c856fc0096e9bc

Observation b2ccb11c-209a-402c-ae20-0d9270cbc34b · outbound

This paper cites A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation.

Combinatorial Reinforcement Learning with Preference Feedback A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.024732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.024732Z digest=sha256:d36faaad4525d88d03e82ed96d3224f59ff0e60dc43038dbe86b7c76df631794

Observation 61a64b14-a5d7-4c2d-91d6-8eb528de124a · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Combinatorial Reinforcement Learning with Preference Feedback Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.158264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.054874Z digest=sha256:280d169d179384019c8afba5137870258dbf3706baccea1af2283c0983e756a1

Observation b708c4e8-4c11-4f98-97b7-749ec3e6d313 · outbound

This paper cites Provably efficient reinforcement learning for discounted mdps with feature mapping.

Combinatorial Reinforcement Learning with Preference Feedback Provably efficient reinforcement learning for discounted mdps with feature mapping

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.048206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.080977Z digest=sha256:2efaad2437591c719cf2c5020f858c8bca8d6a1857aedf384e9934f381278a8b

Observation 1d0be506-c88f-4589-a7c3-2ea55a333c73 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Combinatorial Reinforcement Learning with Preference Feedback Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:13.941327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.108373Z digest=sha256:1b2d40476aab657b380e793f2db9e612d647faff1a728a057c862e84e9f4a8a6

Observation f862ff19-2f46-4f8b-b608-2e8bc716e90a · outbound

This paper cites write newline.

Combinatorial Reinforcement Learning with Preference Feedback write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.158975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.158975Z digest=sha256:847ea9df175fd6698017a9e9b4d5068f48d5b161d7519ac3cef2aea1a2080cef

Pith citing papers

Observation 43b8d225-c795-465e-b2a8-95d9a7baee25 · inbound

Improved Online Confidence Bounds for Multinomial Logistic Bandits cites this paper.

Improved Online Confidence Bounds for Multinomial Logistic Bandits Combinatorial Reinforcement Learning with Preference Feedback

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:57:04.564210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T19:57:04.381612Z digest=sha256:250188b115c4345b2c470b63823842ccb40ca12e8dba958585fcf475fbf2d63f