Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.158975Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 1 inbound Pith citation observation for arXiv:2502.10158.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.158975Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T19:57:04.381612Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T19:57:04.556666Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b18648c4-8d23-482a-b0d4-776f4aea6e69 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Instance-wise minimax-optimal algorithms for logistic bandits
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28165428-c5d9-4984-b1fe-ef7bca76f6de · outbound
Combinatorial Reinforcement Learning with Preference Feedback Vo q l: Towards optimal regret in model-free rl with nonlinear function approximation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f556e53-d9b4-478a-93e5-6bad9c169184 · outbound
Combinatorial Reinforcement Learning with Preference Feedback A tractable online learning algorithm for the multinomial logit contextual bandit
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 548ffc32-b4db-40e6-afa2-84473a54d7ee · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Goyal, N
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5559935-c99c-4704-94fc-f6e73a4e3652 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Thompson sampling for the mnl-bandit
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0080f2ad-5033-432b-92f4-b3169a7009ab · outbound
Combinatorial Reinforcement Learning with Preference Feedback Mnl-bandit: A dynamic learning approach to assortment selection
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a6b6990-e2f1-4953-b1bb-3f9504f68356 · outbound
Combinatorial Reinforcement Learning with Preference Feedback April: Active preference learning-based reinforcement learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ce715ec-61b0-468f-9f9a-9d88960490eb · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Thrampoulidis, C
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87e0a6b4-94a7-4b9e-af07-8a22e8837bc7 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Distributional off-policy evaluation for slate recommendations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 280622c9-3417-4960-854e-bad58b1406bd · outbound
Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2ba47c7e-b279-4d54-8431-699597d2ea37 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Combinatorial multi-armed bandit: General framework and applications
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f373dd7e-7ebb-408d-b664-499d8c7959c3 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1301a24b-9616-49e4-be44-00e3b0eebef5 · outbound
Combinatorial Reinforcement Learning with Preference Feedback F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f827c34-880e-4eb0-b828-1d6e72e67e0d · outbound
Combinatorial Reinforcement Learning with Preference Feedback S., Proutiere, A., et al
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f486ebc0-5827-49a4-b03e-e77ade3febf1 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6efd8f5d-8ecd-4f79-a205-07b39d9a6b00 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Assortment planning under the multinomial logit model with totally unimodular constraint structures
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d28026fd-a4ba-4d6e-99d7-21851d4eb11f · outbound
Combinatorial Reinforcement Learning with Preference Feedback Reinforcement learning with combinatorial actions: An application to vehicle routing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dfac6861-d89d-44c1-9e18-c0cfc3117ede · outbound
Combinatorial Reinforcement Learning with Preference Feedback Bilinear classes: A structural framework for provable generalization in rl
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac33eec9-1542-41d9-aa9a-81687450b21a · outbound
Combinatorial Reinforcement Learning with Preference Feedback Cascading Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfa91a6f-ffe0-4abc-946c-00885a286f88 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Improved optimistic algorithms for logistic bandits
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 098c220a-6abd-4050-b907-6c19d2a35cb0 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Jointly efficient and optimal algorithms for logistic bandits
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e340fbd-91da-4e35-ba4b-b239c9866093 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Parametric bandits: The generalized linear case
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f8a65c8-fb8b-4f3f-8849-0f9032f82e28 · outbound
Combinatorial Reinforcement Learning with Preference Feedback The Statistical Complexity of Interactive Decision Making
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03b3cd5-d973-41bd-a51c-8dfa139917cd · outbound
Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6928cb0d-781e-413a-8e6e-441d96001d7d · outbound
Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit Threads
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d67be2f-445b-4ee9-8b72-fb62b692493b · outbound
Combinatorial Reinforcement Learning with Preference Feedback Slateq: A tractable decomposition for reinforcement learning with recommendation sets
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db054e73-4680-4d27-87cf-20070c8075d5 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Randomized exploration in reinforcement learning with general value function approximation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d37ac1b4-fda5-476f-a172-b224f7503961 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb345469-8408-4da7-ab13-a4e73fec3f43 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 768a359f-e859-454d-8b5d-64a5cc65c0f9 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20505fb5-e7b5-4bb6-bba8-f729ebe59dd9 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c7559e90-d365-4d5d-887e-5c33a19d1edf · outbound
Combinatorial Reinforcement Learning with Preference Feedback Online Sub-Sampling for Reinforcement Learning with General Function Approximation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e51f6d31-7122-487c-b078-3f5864f46902 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Cascading bandits: Learning to rank in the cascade model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97c5c92f-6b79-4add-a98d-4cb47ffe6888 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Combinatorial cascading bandits
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a616e48a-f068-49d1-b623-93d9fbe07a2e · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Hutter, M
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d65c601-5eb0-47a9-a4f8-b27ae658b390 · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28b636af-c5f1-4f43-896a-d20548e80e13 · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Oh, M.-h
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aabc5773-222e-482b-97a6-cc8434b97888 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Online learning to rank with features
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7da764-f091-4cd8-8af8-a1b6b9e12d03 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Modelling the choice of residential location
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ea4aeb3-84a7-439d-8dd2-8f0fb9d3985f · outbound
Combinatorial Reinforcement Learning with Preference Feedback Counterfactual evaluation of slate recommendations with sequential reward interactions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fcfff1cd-f4ab-4158-ae4a-b06c3f79319c · outbound
Combinatorial Reinforcement Learning with Preference Feedback Discrete Sequential Prediction of Continuous Actions for Deep RL
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28ccee3-682d-4983-af4c-fc167647d0c5 · outbound
Combinatorial Reinforcement Learning with Preference Feedback M., and Van Erven, T
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56c2cf88-add8-4503-afae-f42231ec6f72 · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 817c122c-11e5-4d5e-a8d0-f173f9c3e92f · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Iyengar, G
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5709ed03-ef9c-4ebf-8579-29b622ec9063 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Online Learning: A Modern Introduction Using Convex Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71557d63-5f85-40df-aeba-389f1c2d1a42 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Training language models to follow instructions with human feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44bb0a4-8212-49ba-821f-f13f5e3bc851 · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Goyal, V
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a975f62f-39bd-43ad-aa70-cfa62807d193 · outbound
Combinatorial Reinforcement Learning with Preference Feedback M., and Shmoys, D
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fd73fb0f-6618-47e3-8b24-2fa8b5fc3fcd · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Van Roy, B
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3b9cb725-dbc2-4f30-886e-b86fde07259b · outbound
Combinatorial Reinforcement Learning with Preference Feedback CAQL: Continuous Action Q-Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c361c8e7-8103-4598-9c9c-54a000858ecd · outbound
Combinatorial Reinforcement Learning with Preference Feedback Dueling rl: Reinforcement learning with trajectory preferences
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f14d212-1104-4eb3-a749-0100b3ac460c · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Zeevi, A
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e54c55dc-8a4a-4da0-b233-80cb91fa2068 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Deep Reinforcement Learning with Attention for Slate Markov Decision Processes with High-Dimensional States and Actions
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6c0153-9fd0-48af-aeb0-753d3caa5d61 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Off-policy evaluation for slate recommendation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 966e96de-4d75-4ed4-a76e-7b8f27233eb2 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Composite convex minimization involving self-concordant-like cost functions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0381e4e0-bdaa-41eb-8fd2-05ade2ab0bc1 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Control variates for slate off-policy evaluation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94a32340-6a5b-4085-bb4e-f7ca283eb595 · outbound
Combinatorial Reinforcement Learning with Preference Feedback R., and Yang, L
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf765a08-75ae-4f55-8dba-ffcf813dc963 · outbound
Combinatorial Reinforcement Learning with Preference Feedback S., and Krishnamurthy, A
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f7cfa02b-79b3-4be3-9004-935ad45b2f17 · outbound
Combinatorial Reinforcement Learning with Preference Feedback A survey of preference-based reinforcement learning methods
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 709c594c-281b-4b1f-b016-6bb1a69643b8 · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Wang, M
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba8ba420-fd18-4f34-b324-9b965813f310 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 304bad2e-afe8-4a0c-86ac-823bd00b71a0 · outbound
Combinatorial Reinforcement Learning with Preference Feedback and Sugiyama, M
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b2ccb11c-209a-402c-ae20-0d9270cbc34b · outbound
Combinatorial Reinforcement Learning with Preference Feedback A nearly optimal and low-switching algorithm for reinforcement learning with general function approximation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a64b14-a5d7-4c2d-91d6-8eb528de124a · outbound
Combinatorial Reinforcement Learning with Preference Feedback Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b708c4e8-4c11-4f98-97b7-749ec3e6d313 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Provably efficient reinforcement learning for discounted mdps with feature mapping
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d0be506-c88f-4589-a7c3-2ea55a333c73 · outbound
Combinatorial Reinforcement Learning with Preference Feedback Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f862ff19-2f46-4f8b-b608-2e8bc716e90a · outbound
Combinatorial Reinforcement Learning with Preference Feedback write newline
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b8d225-c795-465e-b2a8-95d9a7baee25 · inbound
Improved Online Confidence Bounds for Multinomial Logistic Bandits Combinatorial Reinforcement Learning with Preference Feedback
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.