Pith. sign in

Paper Citation Record · LEDGER

Combinatorial Reinforcement Learning with Preference Feedback

As of 19 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 1 inbound Pith citation observation for arXiv:2502.10158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.10158 v3

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.158975Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:57:04.381612Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T19:57:04.556666Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact2
  • verified fuzzy42
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b18648c4-8d23-482a-b0d4-776f4aea6e69 · outbound

This paper cites Instance-wise minimax-optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Instance-wise minimax-optimal algorithms for logistic bandits

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.791208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.224384Z digest=sha256:85c0ee0d3b0329c576a906e39903e176d761dc7b9e40c603007f0a2443ae2adc

Observation 28165428-c5d9-4984-b1fe-ef7bca76f6de · outbound

This paper cites Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.761100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.231128Z digest=sha256:52ceb3f4e75ec16dec15112b42eb3dcb8fb0567100ed74b08f449dfd3a2b8685

Observation 2f556e53-d9b4-478a-93e5-6bad9c169184 · outbound

This paper cites A tractable online learning algorithm for the multinomial logit contextual bandit.

Combinatorial Reinforcement Learning with Preference Feedback A tractable online learning algorithm for the multinomial logit contextual bandit

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.731998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.237577Z digest=sha256:2dc00ef26d5220bb094eb6e045c43a0e0c0c4e8d3edace1e0eaabd6693aaf209

Observation 548ffc32-b4db-40e6-afa2-84473a54d7ee · outbound

This paper cites and Goyal, N.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, N

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.706758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.245239Z digest=sha256:6b579ecc9922a983b491f055ea0f39c0289bd5a1b2db2775eadc19b2c34b1241

Observation c5559935-c99c-4704-94fc-f6e73a4e3652 · outbound

This paper cites Thompson sampling for the mnl-bandit.

Combinatorial Reinforcement Learning with Preference Feedback Thompson sampling for the mnl-bandit

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.681536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.256779Z digest=sha256:637dcdde6f4c56f1c97fa5859fafe713406952018bb84aba2dc52e0c4447c5ed

Observation 0080f2ad-5033-432b-92f4-b3169a7009ab · outbound

This paper cites Mnl-bandit: A dynamic learning approach to assortment selection.

Combinatorial Reinforcement Learning with Preference Feedback Mnl-bandit: A dynamic learning approach to assortment selection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.657502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.262831Z digest=sha256:3e1539a8f31d1c13a7728d284cb3e593f89406b4bc6b3ae1cbdd204fde37284c

Observation 1a6b6990-e2f1-4953-b1bb-3f9504f68356 · outbound

This paper cites April: Active preference learning-based reinforcement learning.

Combinatorial Reinforcement Learning with Preference Feedback April: Active preference learning-based reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.271517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.271517Z digest=sha256:d9d7b15fb488aae7f2d2d05c5c6c8de7304107ac70bf1353fcdc38bc7f910f1a

Observation 6ce715ec-61b0-468f-9f9a-9d88960490eb · outbound

This paper cites and Thrampoulidis, C.

Combinatorial Reinforcement Learning with Preference Feedback and Thrampoulidis, C

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.611297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.284760Z digest=sha256:ca2e301563d5020a9a3ef1c943c23de651d102d300efc512d856b49d1305459f

Observation 87e0a6b4-94a7-4b9e-af07-8a22e8837bc7 · outbound

This paper cites Distributional off-policy evaluation for slate recommendations.

Combinatorial Reinforcement Learning with Preference Feedback Distributional off-policy evaluation for slate recommendations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.587861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.292863Z digest=sha256:0569c8aaf4070761b58e02af0954457f182e92d8592bc7d81a631ca5029aceb1

Observation 280622c9-3417-4960-854e-bad58b1406bd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.561019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.304989Z digest=sha256:92cb2f63b8e7c821e9c491a0468164c5bb2f1b4b4d69c93b2f74eaee97267d4b

Observation 2ba47c7e-b279-4d54-8431-699597d2ea37 · outbound

This paper cites Combinatorial multi-armed bandit: General framework and applications.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial multi-armed bandit: General framework and applications

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.532539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.312614Z digest=sha256:8c3613d0d97e1bd7956aecd77f34614f08655311d834dd14dc26ca3c21ce1afe

Observation f373dd7e-7ebb-408d-b664-499d8c7959c3 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.507919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.323399Z digest=sha256:70f46e90ead537e7e467b64f7a0b0ef58ebd4a87a2764e1adbc8c9cc9e23dbf3

Observation 1301a24b-9616-49e4-be44-00e3b0eebef5 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Combinatorial Reinforcement Learning with Preference Feedback F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.332237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.332237Z digest=sha256:e4638be0d85b76348e21a4f30b870c7c2cf5aca72192ba8a939f5e39e81b71ac

Observation 9f827c34-880e-4eb0-b828-1d6e72e67e0d · outbound

This paper cites S., Proutiere, A., et al.

Combinatorial Reinforcement Learning with Preference Feedback S., Proutiere, A., et al

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.443014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.341524Z digest=sha256:410e41ad662e89070239df472882d00b0e92f207709d0716490423aabec0ffaf

Observation f486ebc0-5827-49a4-b03e-e77ade3febf1 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.415420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.352173Z digest=sha256:fd50242e9b65f1d785eaa61ca9aff064d28c05e8603071e74cc533cabd9a7a34

Observation 6efd8f5d-8ecd-4f79-a205-07b39d9a6b00 · outbound

This paper cites Assortment planning under the multinomial logit model with totally unimodular constraint structures.

Combinatorial Reinforcement Learning with Preference Feedback Assortment planning under the multinomial logit model with totally unimodular constraint structures

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.383925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.359505Z digest=sha256:af2c11294251e6ac442d9d317ee6493934f75677303f4d772f182f788be00351

Observation d28026fd-a4ba-4d6e-99d7-21851d4eb11f · outbound

This paper cites Reinforcement learning with combinatorial actions: An application to vehicle routing.

Combinatorial Reinforcement Learning with Preference Feedback Reinforcement learning with combinatorial actions: An application to vehicle routing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.337692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.365124Z digest=sha256:feb91351e01696b0cdec49553220f2093cd771414d5d3ca43e3e84e86c3f738b

Observation dfac6861-d89d-44c1-9e18-c0cfc3117ede · outbound

This paper cites Bilinear classes: A structural framework for provable generalization in rl.

Combinatorial Reinforcement Learning with Preference Feedback Bilinear classes: A structural framework for provable generalization in rl

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.314252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.371484Z digest=sha256:62fc3300b2f4d914880ce7654385e56d692ab39492b6e7b538e6fdbc4fa744f4

Observation ac33eec9-1542-41d9-aa9a-81687450b21a · outbound

This paper cites Cascading Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Cascading Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.850736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.377776Z digest=sha256:9a3055bd9df04209dca7695b5a26571ecbb28d4708258a13f10589feb99076fa

Observation cfa91a6f-ffe0-4abc-946c-00885a286f88 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Improved optimistic algorithms for logistic bandits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.405443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.405443Z digest=sha256:de8b3d4eecef83c74cbe4e8e3e3468ab2f6a119976c2aae8d9681e6296159bae

Observation 098c220a-6abd-4050-b907-6c19d2a35cb0 · outbound

This paper cites Jointly efficient and optimal algorithms for logistic bandits.

Combinatorial Reinforcement Learning with Preference Feedback Jointly efficient and optimal algorithms for logistic bandits

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.269061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.475614Z digest=sha256:fe8b525d6844d5acc2ed7b15f60c2c1805023e35a16edab1ab0cc81e46a4aa03

Observation 7e340fbd-91da-4e35-ba4b-b239c9866093 · outbound

This paper cites Parametric bandits: The generalized linear case.

Combinatorial Reinforcement Learning with Preference Feedback Parametric bandits: The generalized linear case

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.243063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.546882Z digest=sha256:b9440159e31f1701a544f29e7bb6861249f891ebbada455e61c1e2ec130f712e

Observation 2f8a65c8-fb8b-4f3f-8849-0f9032f82e28 · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Combinatorial Reinforcement Learning with Preference Feedback The Statistical Complexity of Interactive Decision Making

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.598167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.598167Z digest=sha256:657ac443b073685b562828869ed92c9b296a9e3f42bcb1bf417de3810ff4b4d7

Observation c03b3cd5-d973-41bd-a51c-8dfa139917cd · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.218068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.655974Z digest=sha256:b51dea607a5c08a5ab96525f48db1645899dadf1d837c8e75904c45e4aebfdee

Observation 6928cb0d-781e-413a-8e6e-441d96001d7d · outbound

This paper cites Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:20:13.748595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.665758Z digest=sha256:374465634813c3616a15244eae069790e715b158880b2ef6acab9e6f4c122008

Observation 3d67be2f-445b-4ee9-8b72-fb62b692493b · outbound

This paper cites Slateq: A tractable decomposition for reinforcement learning with recommendation sets.

Combinatorial Reinforcement Learning with Preference Feedback Slateq: A tractable decomposition for reinforcement learning with recommendation sets

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.197624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.672769Z digest=sha256:ee20703a6f62c1d82256aa8d0d2269058397dffa7f5504143f137e9c1b247d21

Observation db054e73-4680-4d27-87cf-20070c8075d5 · outbound

This paper cites Randomized exploration in reinforcement learning with general value function approximation.

Combinatorial Reinforcement Learning with Preference Feedback Randomized exploration in reinforcement learning with general value function approximation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.680375Z digest=sha256:78b7e49479ce5cf829a805a83b6c428296292321b979c19d494c5f3370d23c5f

Observation d37ac1b4-fda5-476f-a172-b224f7503961 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.133885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.685870Z digest=sha256:7b12a78b2736462809860fe3b8fe7b84c98d047aaa1a305bb1173a3d2d08b851

Observation eb345469-8408-4da7-ab13-a4e73fec3f43 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.102859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.692642Z digest=sha256:d6aff74bb1ad855e36237016ea6344a396570b733aaaaf6272380887d982231b

Observation 768a359f-e859-454d-8b5d-64a5cc65c0f9 · outbound

This paper cites an unresolved cited work.

Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T19:20:15.076408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.699888Z digest=sha256:fe6e5329626381391327b7f829fa51b58b424126bbf12da20893ce3484139175

Observation 20505fb5-e7b5-4bb6-bba8-f729ebe59dd9 · outbound

This paper cites Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms.

Combinatorial Reinforcement Learning with Preference Feedback Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.053824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.705775Z digest=sha256:cf41642ff14025f9b35133b149b49d9b2204c2b01bca1ae7f01d396208af7a3a

Observation c7559e90-d365-4d5d-887e-5c33a19d1edf · outbound

This paper cites Online Sub-Sampling for Reinforcement Learning with General Function Approximation.

Combinatorial Reinforcement Learning with Preference Feedback Online Sub-Sampling for Reinforcement Learning with General Function Approximation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.711779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.711779Z digest=sha256:0bbc08621ba1a7356d29cdd964821bb1f619ae2691b246ada243efe4a9fc699b

Observation e51f6d31-7122-487c-b078-3f5864f46902 · outbound

This paper cites Cascading bandits: Learning to rank in the cascade model.

Combinatorial Reinforcement Learning with Preference Feedback Cascading bandits: Learning to rank in the cascade model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:15.022919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.717594Z digest=sha256:2631fb3174bd65913003d889fdbb05821b9d897d77d999629c91ba249604fc47

Observation 97c5c92f-6b79-4add-a98d-4cb47ffe6888 · outbound

This paper cites Combinatorial cascading bandits.

Combinatorial Reinforcement Learning with Preference Feedback Combinatorial cascading bandits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.990933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.722642Z digest=sha256:71e735c7bfb8644744bd35c92f4b2a780d35ebdd3e4775953c1a42688a1a8ad8

Observation a616e48a-f068-49d1-b623-93d9fbe07a2e · outbound

This paper cites and Hutter, M.

Combinatorial Reinforcement Learning with Preference Feedback and Hutter, M

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.969336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.727018Z digest=sha256:7fde2354f1f8e34db0a39434b07f8abe0c515712c50b0e60718ffae1af6e6f77

Observation 1d65c601-5eb0-47a9-a4f8-b27ae658b390 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.946498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.732283Z digest=sha256:97379e40e35f13992ab857495ffc4729c8fe76a5d9dcbdfa98eb84ff98ecb29f

Observation 28b636af-c5f1-4f43-896a-d20548e80e13 · outbound

This paper cites and Oh, M.-h.

Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.921457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.736731Z digest=sha256:0ea1e9cffca85164754091f6767491bc9eaab06eeddd80a4acc258765bddebb9

Observation aabc5773-222e-482b-97a6-cc8434b97888 · outbound

This paper cites Online learning to rank with features.

Combinatorial Reinforcement Learning with Preference Feedback Online learning to rank with features

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.741426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.741426Z digest=sha256:e053d9bed67b1a55d0aa97a0a644d7150ae3235025296d6ca8be1ee3bde2b093

Observation 2e7da764-f091-4cd8-8af8-a1b6b9e12d03 · outbound

This paper cites Modelling the choice of residential location.

Combinatorial Reinforcement Learning with Preference Feedback Modelling the choice of residential location

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.869998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.747449Z digest=sha256:9b57c0b24293635aba37eaab358cd8b52769d9f460bfe967f60f55cba928ec31

Observation 7ea4aeb3-84a7-439d-8dd2-8f0fb9d3985f · outbound

This paper cites Counterfactual evaluation of slate recommendations with sequential reward interactions.

Combinatorial Reinforcement Learning with Preference Feedback Counterfactual evaluation of slate recommendations with sequential reward interactions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.846788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.754104Z digest=sha256:3c412a395c493e4014682ad2702179d2fdce8ce8f73ddcdd9914f7d922134998

Observation fcfff1cd-f4ab-4158-ae4a-b06c3f79319c · outbound

This paper cites Discrete Sequential Prediction of Continuous Actions for Deep RL.

Combinatorial Reinforcement Learning with Preference Feedback Discrete Sequential Prediction of Continuous Actions for Deep RL

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.761945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.761945Z digest=sha256:18fb7cd6b7be6f580f475e9ee27325ba484c757ff23d5a234cfc15469cddfcbf

Observation f28ccee3-682d-4983-af4c-fc167647d0c5 · outbound

This paper cites M., and Van Erven, T.

Combinatorial Reinforcement Learning with Preference Feedback M., and Van Erven, T

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.823133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.792670Z digest=sha256:e7dbc458043d939a8fa9e0c46ac5ab1e08900bbeb189a4239169b1ca0579a73a

Observation 56c2cf88-add8-4503-afae-f42231ec6f72 · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.798459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.825505Z digest=sha256:d15a74114ad2c0b0b0940666b7fb7b089902b1d07106cae5e7c97c3df9d1394c

Observation 817c122c-11e5-4d5e-a8d0-f173f9c3e92f · outbound

This paper cites and Iyengar, G.

Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.776371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.844502Z digest=sha256:53ad02e55acb040f0af340763fd7fd8c504791a0e7853172273aaf12c17fe8ce

Observation 5709ed03-ef9c-4ebf-8579-29b622ec9063 · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Combinatorial Reinforcement Learning with Preference Feedback Online Learning: A Modern Introduction Using Convex Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.870065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.870065Z digest=sha256:39fef138d6d6f3157966339b6292df0fb8d971d86733c37f6a8d58f8a09bcb15

Observation 71557d63-5f85-40df-aeba-389f1c2d1a42 · outbound

This paper cites Training language models to follow instructions with human feedback.

Combinatorial Reinforcement Learning with Preference Feedback Training language models to follow instructions with human feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.896163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.896163Z digest=sha256:fccb500b77f03eb51425a7cd972e87fb7222ecb1fc2c69f7b5563e134de5d530

Observation b44bb0a4-8212-49ba-821f-f13f5e3bc851 · outbound

This paper cites and Goyal, V.

Combinatorial Reinforcement Learning with Preference Feedback and Goyal, V

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.725451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.904201Z digest=sha256:6d8368acade4a29995a70ce6720f979255b1ad0af6ea43ea5cf9617d53aedf2d

Observation a975f62f-39bd-43ad-aa70-cfa62807d193 · outbound

This paper cites M., and Shmoys, D.

Combinatorial Reinforcement Learning with Preference Feedback M., and Shmoys, D

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.702565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.910945Z digest=sha256:90f9eeafd645b9dcd1d1d4bc581f8b4c5b580824c5139092f6989867c1850efe

Observation fd73fb0f-6618-47e3-8b24-2fa8b5fc3fcd · outbound

This paper cites and Van Roy, B.

Combinatorial Reinforcement Learning with Preference Feedback and Van Roy, B

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.674460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.918859Z digest=sha256:89aab103c344cbf0b8f77654044efe53526a2eac61281fa674a93fe6c6d47c4e

Observation 3b9cb725-dbc2-4f30-886e-b86fde07259b · outbound

This paper cites CAQL: Continuous Action Q-Learning.

Combinatorial Reinforcement Learning with Preference Feedback CAQL: Continuous Action Q-Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.926888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.926888Z digest=sha256:35683f326c0bd5433d7d0cc4cd88f86347600b35603f92daf715bcda09d6388b

Observation c361c8e7-8103-4598-9c9c-54a000858ecd · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Combinatorial Reinforcement Learning with Preference Feedback Dueling rl: Reinforcement learning with trajectory preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.934451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.934451Z digest=sha256:792d3bf570e01f433fba387224fe0c88a51e70cc424c97870fba30a47c50983b

Observation 0f14d212-1104-4eb3-a749-0100b3ac460c · outbound

This paper cites and Zeevi, A.

Combinatorial Reinforcement Learning with Preference Feedback and Zeevi, A

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.571717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.943188Z digest=sha256:36ad399f7e97061ccf6c1f045a921543cc03ed2ef26de4c6da9983418d57035d

Observation e54c55dc-8a4a-4da0-b233-80cb91fa2068 · outbound

This paper cites Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions.

Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.951927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.951927Z digest=sha256:8c4a29e36e5f35184b4a7d1cf9cd00069bd52812e44e7880efc6a4bd3e4f6efa

Observation 2d6c0153-9fd0-48af-aeb0-753d3caa5d61 · outbound

This paper cites Off-policy evaluation for slate recommendation.

Combinatorial Reinforcement Learning with Preference Feedback Off-policy evaluation for slate recommendation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.490085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.958603Z digest=sha256:f58e3e52080d40a11333be5b1790d8b47633efac43c5e5c08e2e657c393888d1

Observation 966e96de-4d75-4ed4-a76e-7b8f27233eb2 · outbound

This paper cites Composite convex minimization involving self-concordant-like cost functions.

Combinatorial Reinforcement Learning with Preference Feedback Composite convex minimization involving self-concordant-like cost functions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.392696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.967403Z digest=sha256:f87c6261acd670d90c63440cb516332f4f5a0bf77c5e4ba6f014ac97d88f8d21

Observation 0381e4e0-bdaa-41eb-8fd2-05ade2ab0bc1 · outbound

This paper cites Control variates for slate off-policy evaluation.

Combinatorial Reinforcement Learning with Preference Feedback Control variates for slate off-policy evaluation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.370807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.973733Z digest=sha256:c114adb5c024560974b74d66484b80d5a1b0aead39925f894ebc166c2a39aa95

Observation 94a32340-6a5b-4085-bb4e-f7ca283eb595 · outbound

This paper cites R., and Yang, L.

Combinatorial Reinforcement Learning with Preference Feedback R., and Yang, L

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.343071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.979355Z digest=sha256:6591b30e58caab3c1f756c6e4e49b8a80f191ddd95a262f6a2e7b5afe3b16e93

Observation cf765a08-75ae-4f55-8dba-ffcf813dc963 · outbound

This paper cites S., and Krishnamurthy, A.

Combinatorial Reinforcement Learning with Preference Feedback S., and Krishnamurthy, A

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.299861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.984567Z digest=sha256:c9e608070624d4c5e7f74081668b0d8839a0519d9a72f80e078e96c9eca5413e

Observation f7cfa02b-79b3-4be3-9004-935ad45b2f17 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Combinatorial Reinforcement Learning with Preference Feedback A survey of preference-based reinforcement learning methods

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:12.992315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:12.992315Z digest=sha256:1f0a79200af6b351a7e360fef9a8ee9228a4a43a1fd3d02544dd525ff0ec10b0

Observation 709c594c-281b-4b1f-b016-6bb1a69643b8 · outbound

This paper cites and Wang, M.

Combinatorial Reinforcement Learning with Preference Feedback and Wang, M

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.250244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:12.998266Z digest=sha256:a9f9b409a56eca97a08baaef50d792938305fc639b1828a5824a15af88f52918

Observation ba8ba420-fd18-4f34-b324-9b965813f310 · outbound

This paper cites Provable Offline Preference-Based Reinforcement Learning.

Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.002840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.002840Z digest=sha256:b37dca615461f691fb54cd5bb6cc2f6551c72dcb98e84dc43fb41a32ee64edde

Observation 304bad2e-afe8-4a0c-86ac-823bd00b71a0 · outbound

This paper cites and Sugiyama, M.

Combinatorial Reinforcement Learning with Preference Feedback and Sugiyama, M

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.224992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.009631Z digest=sha256:c6b88ff279fb4dc5e862f81c102ff281c9159afa4ba43e15f6ab62c6f31ae83c

Observation b2ccb11c-209a-402c-ae20-0d9270cbc34b · outbound

This paper cites A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation.

Combinatorial Reinforcement Learning with Preference Feedback A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.024732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.024732Z digest=sha256:5ff3b426d4b0d6602000cc79e152a85744cad9557fd8f50f7fa7b8d723fbf9fb

Observation 61a64b14-a5d7-4c2d-91d6-8eb528de124a · outbound

This paper cites Nearly minimax optimal reinforcement learning for linear mixture markov decision processes.

Combinatorial Reinforcement Learning with Preference Feedback Nearly minimax optimal reinforcement learning for linear mixture markov decision processes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.158264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.054874Z digest=sha256:3fa00db2e152afd03d0c0b1cafef9e5573623e96ccb9d77ad7bb27f793d0adc3

Observation b708c4e8-4c11-4f98-97b7-749ec3e6d313 · outbound

This paper cites Provably efficient reinforcement learning for discounted mdps with feature mapping.

Combinatorial Reinforcement Learning with Preference Feedback Provably efficient reinforcement learning for discounted mdps with feature mapping

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:14.048206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.080977Z digest=sha256:33d4ff6a7d2350dc4bcb72c4987d1397c1607dda2cf645ae3cdfc4a6759e51d2

Observation 1d0be506-c88f-4589-a7c3-2ea55a333c73 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Combinatorial Reinforcement Learning with Preference Feedback Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T19:20:13.941327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:20:13.108373Z digest=sha256:66d849a9c280c633c31e721b520c0e64dc90373afffb2f219ad20ee2215d7d00

Observation f862ff19-2f46-4f8b-b608-2e8bc716e90a · outbound

This paper cites write newline.

Combinatorial Reinforcement Learning with Preference Feedback write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.158975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.158975Z digest=sha256:9eb6b27c699cfa456aa08f5abb1b4e3123b88c368107043ba3c0ebb2137a2ae9

Pith citing papers

Observation 43b8d225-c795-465e-b2a8-95d9a7baee25 · inbound

Improved Online Confidence Bounds for Multinomial Logistic Bandits cites this paper.

Improved Online Confidence Bounds for Multinomial Logistic Bandits Combinatorial Reinforcement Learning with Preference Feedback

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T19:57:04.564210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T19:57:04.381612Z digest=sha256:63d1014c45217c70cb9b556be1356e0b513650ac2e16c434f6cf672251e52111