Pith. sign in

Paper Citation Record · LEDGER

Contextual bandits with entropy-based human feedback

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2502.08759.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08759 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:51:35.789930Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T05:06:05.369793Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.303918Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b329e6c-839b-45aa-be4a-452e93697d19 · outbound

This paper cites GPT-4 Technical Report.

Contextual bandits with entropy-based human feedback GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.699578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.699578Z digest=sha256:16a6b882f173fbd8e22f3aae323e4d29551d9e30f7396a67b80364d3fc8a5525

Observation f42ea444-bab5-46ff-b8b5-c3776a0f9f52 · outbound

This paper cites Survey on appli- cations of multi-armed and contextual bandits.

Contextual bandits with entropy-based human feedback Survey on appli- cations of multi-armed and contextual bandits

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.257362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.728588Z digest=sha256:4cade7807c6ddee0a16e3a89256b6679577b38863ac0a9a35b55506a059f83b4

Observation 043df074-a09e-41f9-ade2-37d263d88070 · outbound

This paper cites Thompson sampling with the online bootstrap.

Contextual bandits with entropy-based human feedback Thompson sampling with the online bootstrap

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.746581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.746581Z digest=sha256:8bc349359d26f9bf0f367a6ec08cd965dfaf52b97cb9c7a4c8b8f308d6e149f0

Observation 60c0043b-797e-44af-929d-60620289ba07 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Contextual bandits with entropy-based human feedback Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.756299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.756299Z digest=sha256:52c3ab21b6c48b2206d21bdb6e00a77d65ed0c7e6d55fb6d6fc23f77b2d1ea6f

Observation 6f654b1b-2411-418f-976b-9e2e4d522bcc · outbound

This paper cites and Lefebvre, S.

Contextual bandits with entropy-based human feedback and Lefebvre, S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.227895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.765013Z digest=sha256:042a415b0d2d9de91ac20b5abd6ad5df8d15716576f707cfe9ec650a64fd77cf

Observation 9a2641ea-dd53-4d65-83f8-4d8838d5908f · outbound

This paper cites Borda Regret Minimization for Generalized Linear Dueling Bandits.

Contextual bandits with entropy-based human feedback Borda Regret Minimization for Generalized Linear Dueling Bandits

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:51:35.866329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.769075Z digest=sha256:855104aff478120729002333b8893d226c65df01b94e35330ee8a7b6d1db5f3f

Observation 3eb569e2-74e1-4e2e-b5be-a6537b2ea515 · outbound

This paper cites FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback.

Contextual bandits with entropy-based human feedback FRESH: Interactive Reward Shaping in High-Dimensional State Spaces using Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.773348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.773348Z digest=sha256:904c497a767d792f156bbd93957b70598bf1257fb4077ff77f999bbc478ae63b

Observation 68945443-7639-49d9-a6f4-53eb5da189de · outbound

This paper cites CAREForMe: Contextual Multi-Armed Bandit Recommendation Framework for Mental Health.

Contextual bandits with entropy-based human feedback CAREForMe: Contextual Multi-Armed Bandit Recommendation Framework for Mental Health

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:51:35.829846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.778358Z digest=sha256:e301e60e880dcfbbbd808bfa2b8e70bc9226f60ba3d9f97a62af86e8daaac42b

Observation b027f9b1-4418-4521-88de-50160e7142ea · outbound

This paper cites These connections highlight how our approach advances real-time feedback integration and decision optimization.

Contextual bandits with entropy-based human feedback These connections highlight how our approach advances real-time feedback integration and decision optimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.185489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.789930Z digest=sha256:48105b3c1f4215e09eb42c2627861042cb7b6c041d0c9d7d1a04ced9b88325e3

Observation 0f97e972-8758-4f94-8edd-8bf40a2ae367 · outbound

This paper cites EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits.

Contextual bandits with entropy-based human feedback EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.714690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.714690Z digest=sha256:c3297e900b74ecb7c14b91ef4a3df18b44ae64bfa8790e9799a87ca5c092de3b

Observation e5f631b0-48e9-4dd1-8bde-3ecdd832eefd · outbound

This paper cites Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism.

Contextual bandits with entropy-based human feedback Reinforcement Learning with Human Feedback: Learning Dynamic Choices via Pessimism

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.751446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.751446Z digest=sha256:49ed6cd03e9bea2dae2969ace797d67cfd49f5574a9f5f3f5553deed0e36c343

Observation b1fc2aaa-78a2-40a4-ad5e-db4c9de09e94 · outbound

This paper cites an unresolved cited work.

Contextual bandits with entropy-based human feedback Unresolved cited work

Reference 2011

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:51:36.199877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.786303Z digest=sha256:558ddf8573f6ded255109c4a32bb0bedee5026b5d117ee4402b16a6305872033

Observation 924ee04e-f949-43f0-befb-3c9a2f7be854 · outbound

This paper cites A neural networks committee for the contextual bandit problem.

Contextual bandits with entropy-based human feedback A neural networks committee for the contextual bandit problem

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.286247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.704685Z digest=sha256:a2b826467d2812df478a0671cd8bdb47b4a0e2a505400bc3da320ccc64968f24

Observation 7c2b461e-b3f5-4629-ad64-d5ed89b25720 · outbound

This paper cites DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback.

Contextual bandits with entropy-based human feedback DQN-TAMER: Human-in-the-Loop Reinforcement Learning with Intractable Feedback

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.709372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.709372Z digest=sha256:d76751b5377f682bb6d753a4f13572ae929998ed3464e6f4aaeb9335584f8823

Observation dcd5ba6d-5710-4445-935e-463beb14eb02 · outbound

This paper cites Bayesian Active Learning for Classification and Preference Learning.

Contextual bandits with entropy-based human feedback Bayesian Active Learning for Classification and Preference Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.741570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.741570Z digest=sha256:f866ecdb0558ad1da68f4da4c172f5aab4def6117576d22557c84cb8b99c4c4c

Observation da1357dc-0c76-444e-a6f2-231d2566678c · outbound

This paper cites Contextual Bandits and Imitation Learning via Preference-Based Active Queries.

Contextual bandits with entropy-based human feedback Contextual Bandits and Imitation Learning via Preference-Based Active Queries

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.760624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.760624Z digest=sha256:637cb4dff61a5d75ecdbf4c1a1289b3df2057ac7777c3a38476e1bd93a016b29

Observation 1463b229-62a1-4005-a5c0-993563767aef · outbound

This paper cites an unresolved cited work.

Contextual bandits with entropy-based human feedback Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:51:36.241736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.732765Z digest=sha256:244b91b86e0cc03ce765c2d1b2bb03a5f2e993160f591afa53859fdb0e80e826

Observation 8f0b24d6-f70e-4d5f-bc2c-b702d3fc6260 · outbound

This paper cites We draw inspiration from Tang and Wiens (Tang & Wiens, 2023), whose counterfactual-augmented importance sampling informs our feedback framework, and extend DAGGER (Ross et al.,.

Contextual bandits with entropy-based human feedback We draw inspiration from Tang and Wiens (Tang & Wiens, 2023), whose counterfactual-augmented importance sampling informs our feedback framework, and extend DAGGER (Ross et al.,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.214241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.782304Z digest=sha256:63fff329b8827587766f419e1e65766c84bdb8d8722b5d9175d0533a260b80fd

Observation 8defbab9-48ec-4a59-985a-af04f1c8837e · outbound

This paper cites Adversarial Rewards in Universal Learning for Contextual Bandits.

Contextual bandits with entropy-based human feedback Adversarial Rewards in Universal Learning for Contextual Bandits

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:51:36.125117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.719330Z digest=sha256:6e678833c077ab852bd01cf49d2239b0e385291b026b865c2fc79f75851b4e72

Observation af5bd6da-e7f1-4604-aa21-281231ee939f · outbound

This paper cites Contextual bandit for active learning: Active thompson sampling.

Contextual bandits with entropy-based human feedback Contextual bandit for active learning: Active thompson sampling

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:51:36.272001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:51:35.724103Z digest=sha256:647755b5bae44a912a1c4b8b19b23447e55c105b912f0bd3de998621e8047727

Observation 41d1a8fb-302e-4ddf-8cde-1e20dcaff48e · outbound

This paper cites Nearly optimal algorithms for con- textual dueling bandits from adversarial feedback.

Contextual bandits with entropy-based human feedback Nearly optimal algorithms for con- textual dueling bandits from adversarial feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T23:51:35.737043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:51:35.737043Z digest=sha256:bd932232e37023c33f5e95be5faaab99034430156b38347aaa765cc4f3962ec6

Pith citing papers

Observation 15283c2d-6fbd-4d8f-bf12-b5bff76920e6 · inbound

Autoformalization of Agent Instructions into Policy-as-Code cites this paper.

Autoformalization of Agent Instructions into Policy-as-Code Contextual bandits with entropy-based human feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.305589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T05:06:05.369793Z digest=sha256:8362db02cb0ea7e7443b289f64d68a330cb36289bf93ed3e355da3eeb5bd60d2