Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:51:35.789930Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2502.08759.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:51:35.789930Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T05:06:05.369793Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:39:50.303918Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3b329e6c-839b-45aa-be4a-452e93697d19 · outbound
Contextual bandits with entropy-based human feedback GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f42ea444-bab5-46ff-b8b5-c3776a0f9f52 · outbound
Contextual bandits with entropy-based human feedback Survey on appli- cations of multi-armed and contextual bandits
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 043df074-a09e-41f9-ade2-37d263d88070 · outbound
Contextual bandits with entropy-based human feedback Thompson sampling with the online bootstrap
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c0043b-797e-44af-929d-60620289ba07 · outbound
Contextual bandits with entropy-based human feedback Proximal Policy Optimization Algorithms
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f654b1b-2411-418f-976b-9e2e4d522bcc · outbound
Contextual bandits with entropy-based human feedback and Lefebvre, S
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a2641ea-dd53-4d65-83f8-4d8838d5908f · outbound
Contextual bandits with entropy-based human feedback Borda Regret Minimization for Generalized Linear Dueling Bandits
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3eb569e2-74e1-4e2e-b5be-a6537b2ea515 · outbound
Contextual bandits with entropy-based human feedback FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68945443-7639-49d9-a6f4-53eb5da189de · outbound
Contextual bandits with entropy-based human feedback CAREForMe: Contextual Multi-Armed Bandit Recommendation Framework for Mental Health
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b027f9b1-4418-4521-88de-50160e7142ea · outbound
Contextual bandits with entropy-based human feedback These connections highlight how our approach advances real-time feedback integration and decision optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0f97e972-8758-4f94-8edd-8bf40a2ae367 · outbound
Contextual bandits with entropy-based human feedback EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits
Reference 2002
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f631b0-48e9-4dd1-8bde-3ecdd832eefd · outbound
Contextual bandits with entropy-based human feedback Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1fc2aaa-78a2-40a4-ad5e-db4c9de09e94 · outbound
Contextual bandits with entropy-based human feedback Unresolved cited work
Reference 2011
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 924ee04e-f949-43f0-befb-3c9a2f7be854 · outbound
Contextual bandits with entropy-based human feedback A neural networks committee for the contextual bandit problem
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c2b461e-b3f5-4629-ad64-d5ed89b25720 · outbound
Contextual bandits with entropy-based human feedback DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd5ba6d-5710-4445-935e-463beb14eb02 · outbound
Contextual bandits with entropy-based human feedback Bayesian Active Learning for Classification and Preference Learning
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1357dc-0c76-444e-a6f2-231d2566678c · outbound
Contextual bandits with entropy-based human feedback Contextual Bandits and Imitation Learning via Preference-Based Active Queries
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1463b229-62a1-4005-a5c0-993563767aef · outbound
Contextual bandits with entropy-based human feedback Unresolved cited work
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f0b24d6-f70e-4d5f-bc2c-b702d3fc6260 · outbound
Contextual bandits with entropy-based human feedback We draw inspiration from Tang and Wiens (Tang & Wiens, 2023), whose counterfactual-augmented importance sampling informs our feedback framework, and extend DAGGER (Ross et al.,
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8defbab9-48ec-4a59-985a-af04f1c8837e · outbound
Contextual bandits with entropy-based human feedback Adversarial Rewards in Universal Learning for Contextual Bandits
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af5bd6da-e7f1-4604-aa21-281231ee939f · outbound
Contextual bandits with entropy-based human feedback Contextual bandit for active learning: Active thompson sampling
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 41d1a8fb-302e-4ddf-8cde-1e20dcaff48e · outbound
Contextual bandits with entropy-based human feedback Nearly optimal algorithms for con- textual dueling bandits from adversarial feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15283c2d-6fbd-4d8f-bf12-b5bff76920e6 · inbound
Autoformalization of Agent Instructions into Policy-as-Code Contextual bandits with entropy-based human feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.