Pith. sign in

Paper Citation Record · LEDGER

Fusing Reward and Dueling Feedback in Stochastic Bandits

As of 19 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2504.15812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15812 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:26:28.023566Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T11:34:19.258512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:36:13.482957Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da95bc04-4358-408a-9646-d6b8a7a366cf · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Improved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.912569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.912569Z digest=sha256:07f4d80fb6ee0a8d05a6f3c5755355e1aeaeb48b4a1e1657847d0e68b1edc005

Observation 8641145e-155e-4d28-a5b4-56a996489663 · outbound

This paper cites Reducing dueling bandits to cardinal bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Reducing dueling bandits to cardinal bandits

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.916058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.916058Z digest=sha256:2abc18e23df4cc0df22ff364d208e08e6a0be475a7a60d15d19a0e3b01d421b4

Observation abfb73c0-70b3-4592-9542-49308f27734f · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs.

Fusing Reward and Dueling Feedback in Stochastic Bandits Using confidence bounds for exploitation-exploration trade-offs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.918940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.918940Z digest=sha256:047494f1e3d230862084ac90288c4143c6921ec063f8b3efe4a21994beaa9ce9

Observation a9cb56b2-6511-4aa0-8461-d2b0c0f39069 · outbound

This paper cites and Ortner, R.

Fusing Reward and Dueling Feedback in Stochastic Bandits and Ortner, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.299414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.923021Z digest=sha256:52a475f4fc45f2a567896d13167a866fe36fa3a4fd2f19549554debcae73d5fb

Observation e49611b4-babd-4f79-a02a-1a5cdd7bc54d · outbound

This paper cites an unresolved cited work.

Fusing Reward and Dueling Feedback in Stochastic Bandits Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:26:28.290627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.926123Z digest=sha256:ba6c1a43c698af8d999af5089ba3ecc909370622212c3274db1c8013593bda5f

Observation ec4067e4-2020-4531-9bd5-6526587055cd · outbound

This paper cites Online Learning and Bandits with Queried Hints.

Fusing Reward and Dueling Feedback in Stochastic Bandits Online Learning and Bandits with Queried Hints

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:26:28.093365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.929035Z digest=sha256:90dfd9668407a4780195a717983ee17374ab059cce5d91fdbe02555f413f7c4b

Observation 95da1968-fc23-4f17-a86f-fb08a005ab0e · outbound

This paper cites Bandits games and clustering foundations.

Fusing Reward and Dueling Feedback in Stochastic Bandits Bandits games and clustering foundations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.281657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.932434Z digest=sha256:ae9fc1d0e0951a387edcdb05eec1c1eca7c2baa3b420f3db8fad7e462ccd7d3c

Observation 9ea97d07-b67a-4b45-96a5-5795d56bebe8 · outbound

This paper cites Regret analysis of stochastic and nonstochastic multi-armed bandit problems.

Fusing Reward and Dueling Feedback in Stochastic Bandits Regret analysis of stochastic and nonstochastic multi-armed bandit problems

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.273168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.935386Z digest=sha256:d3deadaab0de9b18f8bac7cb2aad605fab459db175b0650ad1818904a58733b9

Observation 1229c264-8c7a-4c4f-a734-8e226e4fe506 · outbound

This paper cites Kullback-leibler upper confidence bounds for optimal sequential allocation.

Fusing Reward and Dueling Feedback in Stochastic Bandits Kullback-leibler upper confidence bounds for optimal sequential allocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.264473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.938240Z digest=sha256:10e15b318b5402163d9ef59d2ae1b97f2c92a3c66766da6d2988e26b30c36482

Observation e28f4d59-b268-451e-8539-e332c1032733 · outbound

This paper cites Combinatorial multi-armed bandit: General framework and applications.

Fusing Reward and Dueling Feedback in Stochastic Bandits Combinatorial multi-armed bandit: General framework and applications

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.941236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.941236Z digest=sha256:f10e2da54153b7fd46390ce75a51e4816d455b6f845b4145fac7be47e90d4d58

Observation 4778f404-d45c-4f15-9964-453f651d3f38 · outbound

This paper cites Leveraging initial hints for free in stochastic linear bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Leveraging initial hints for free in stochastic linear bandits

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.250470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.944121Z digest=sha256:eb80ddf674f3561ed1140e1d0adf62f3dee4c363324a88c79c086dea724f7303

Observation 1744d793-f794-4568-8967-d4b05fe3a198 · outbound

This paper cites and Takemura, A.

Fusing Reward and Dueling Feedback in Stochastic Bandits and Takemura, A

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.947306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.947306Z digest=sha256:fdf82316d1b039d882691f1f07d2512631f4aa3e76e033fb7c9025f8b69fc69a

Observation badb5ea3-3f00-44ec-a56a-facf1f0a8407 · outbound

This paper cites Provable Benefits of Policy Learning from Human Preferences in Contextual Bandit Problems.

Fusing Reward and Dueling Feedback in Stochastic Bandits Provable Benefits of Policy Learning from Human Preferences in Contextual Bandit Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.949991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.949991Z digest=sha256:15b0d3b69a6617200dbe9647f868ada4e50e03428d1d78d7492effbefedb99b7

Observation b68548a0-716c-4ed6-a949-d7f44193029f · outbound

This paper cites Regret lower bound and optimal algorithm in dueling bandit problem.

Fusing Reward and Dueling Feedback in Stochastic Bandits Regret lower bound and optimal algorithm in dueling bandit problem

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.235378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.953158Z digest=sha256:7bb374d433277627011dc2c3d9935d9376e465b31071ae6ff4767cc95230d4c3

Observation 8abb0548-3f25-491c-9aa6-3e8450b9550f · outbound

This paper cites an unresolved cited work.

Fusing Reward and Dueling Feedback in Stochastic Bandits Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.955799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.955799Z digest=sha256:0e1e680e15d6be3b3615b2f88cc3715cda7b355e26a795c565db136b6b4ebe67

Observation b5951fdd-7891-4fd5-a09a-36bb5d1b7f13 · outbound

This paper cites and Szepesv \'a ri, C.

Fusing Reward and Dueling Feedback in Stochastic Bandits and Szepesv \'a ri, C

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.958731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.958731Z digest=sha256:a0388c8320e712bfacec3f6ea439e7a5a60470c8e2721027ac5632fd50cda43f

Observation 1331cd2d-e465-49dc-8f34-75795c47aad8 · outbound

This paper cites an unresolved cited work.

Fusing Reward and Dueling Feedback in Stochastic Bandits Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.961510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.961510Z digest=sha256:b680b4077751c22900f332650984494449e391b5c4d872788d9e2221a647d946

Observation 385e472a-2f51-4f3b-8361-4a202634539f · outbound

This paper cites FedConPE: Efficient Federated Conversational Bandits with Heterogeneous Clients.

Fusing Reward and Dueling Feedback in Stochastic Bandits FedConPE: Efficient Federated Conversational Bandits with Heterogeneous Clients

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.964233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.964233Z digest=sha256:9d40274275f516668df1c078d51adb78fd18b7f41be141d6f34ac9001e377fb5

Observation c4e6dd19-4977-4228-9fa2-4b70c6d31105 · outbound

This paper cites Predictive bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Predictive bandits

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.209879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.967344Z digest=sha256:f4b3f3990d40fe6c0dd1bad9750b9a1fb55e29fcf6ef117208f6aa6e0e9cdd8a

Observation 6c36c2ed-5cfe-4df9-8317-9a92750278ea · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Fusing Reward and Dueling Feedback in Stochastic Bandits Online Learning: A Modern Introduction Using Convex Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.970502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.970502Z digest=sha256:b18bb1a46318e70b58f66e02d086f95d3342393e15114a769c964b92dc4853ba

Observation a8566238-90c5-4e1c-8d00-83b638a7f7d7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Fusing Reward and Dueling Feedback in Stochastic Bandits Training language models to follow instructions with human feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.974021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.974021Z digest=sha256:4175492c57820d7b815eb23f4adcec2fc2f63a904a905725eafa336be9811a32

Observation a35474b1-5f3b-47b4-a798-54900651ec60 · outbound

This paper cites D., Ermon, S., and Finn, C.

Fusing Reward and Dueling Feedback in Stochastic Bandits D., Ermon, S., and Finn, C

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.976926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.976926Z digest=sha256:001feaebb8c51e6ed1ea230e2477eb170e0a58d91652f1be3408d06f099ea2a9

Observation 27dae902-6484-4326-8333-a136e8eb860b · outbound

This paper cites and Gaillard, P.

Fusing Reward and Dueling Feedback in Stochastic Bandits and Gaillard, P

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.190216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.979839Z digest=sha256:b65d736bbafa362024851d06bb21353c3c3c4f52ad4dc5709f6722ae08fdf00a

Observation df60c913-3808-40a8-bc17-823aeb6631f6 · outbound

This paper cites Adversarial dueling bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Adversarial dueling bandits

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.181819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.982972Z digest=sha256:4106ee4eacc145930f9f5464f415468b304a3855f55abe091c62f1881b1e30f7

Observation cc7c96b8-eb21-4a36-aeff-7036b49e9a89 · outbound

This paper cites Introduction to multi-armed bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Introduction to multi-armed bandits

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.985879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.985879Z digest=sha256:ea406a380904004f1998c88dfb9ec15bb5ba2c407193d884f0cac258809e26a9

Observation 9efd110b-a94a-425a-9fab-213ee64c2fbc · outbound

This paper cites Advancements in dueling bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Advancements in dueling bandits

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.173745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.989861Z digest=sha256:08b204b7be5b841e097942311577b9e59144bafc8c477c031ca34895e915be8e

Observation cde7e502-452d-42db-b33a-a41576798b3e · outbound

This paper cites S., Barto, A.

Fusing Reward and Dueling Feedback in Stochastic Bandits S., Barto, A

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.992859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.992859Z digest=sha256:4c66b7b031de5c80fc44426af9b72811f0eebc888be87e95e05f8fe35f2b75f1

Observation d2af2e41-4d79-47fc-b6a8-91c1df46c467 · outbound

This paper cites Algorithms for reinforcement learning.

Fusing Reward and Dueling Feedback in Stochastic Bandits Algorithms for reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:27.996006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:27.996006Z digest=sha256:4a3c71aa26a24d364b8112b42256bd188ee9b7dd6e80aad79639d889eb1d41ff

Observation f26268ef-941e-46ba-8bfb-5d0e3ed82d7d · outbound

This paper cites Generic exploration and k-armed voting bandits.

Fusing Reward and Dueling Feedback in Stochastic Bandits Generic exploration and k-armed voting bandits

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.155100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:27.999561Z digest=sha256:3c0a5cbeb8d1de634392a0ba7b1abc6aff867c36e34ef23112219c293797d85f

Observation 1889f57a-1754-46e2-aa64-b63558dcbddc · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Fusing Reward and Dueling Feedback in Stochastic Bandits Is RLHF More Difficult than Standard RL?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:28.002503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:28.002503Z digest=sha256:1d541024a77f5e7e70a2da9a27e879de44122b41116a42b623f1b45a37cdcd50

Observation 3e683b71-0839-4829-acaa-773fd1fcf329 · outbound

This paper cites an unresolved cited work.

Fusing Reward and Dueling Feedback in Stochastic Bandits Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:26:28.146612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:28.005792Z digest=sha256:2a492f1459bb6eb63a01fd3ec0c792a0fa035942f18be40146227ad2a9294e16

Observation 66522df1-fb9f-41b7-a3d3-3b2240536954 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Fusing Reward and Dueling Feedback in Stochastic Bandits Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:28.008906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:28.008906Z digest=sha256:105893c6905bae4ddba367d38a7a4622147152e5b726ef08115d5a1c058f5e5d

Observation c0bc00a5-59cf-4316-a1c8-059998ba029d · outbound

This paper cites L., Craswell, N., Voorhees, E.

Fusing Reward and Dueling Feedback in Stochastic Bandits L., Craswell, N., Voorhees, E

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.132593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:28.011819Z digest=sha256:ecc1ddb4c05274ede4e76c794651d271b44d8544e950a2e76584f4e13cf9a071

Observation ce8bcfb3-25fa-448f-b268-22d5b455dfb3 · outbound

This paper cites The k-armed dueling bandits problem.

Fusing Reward and Dueling Feedback in Stochastic Bandits The k-armed dueling bandits problem

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:28.014581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:28.014581Z digest=sha256:371de9b6a7378e150274f53f22ec2febe35eee420e6ba781111e6fcdbdbb331b

Observation 17366d05-a6f8-4780-b2f5-76e1297ab72d · outbound

This paper cites Multi-armed bandit with additional observations.

Fusing Reward and Dueling Feedback in Stochastic Bandits Multi-armed bandit with additional observations

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.118696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:28.017831Z digest=sha256:4b0bcca8cdd7f4cde2e6f7509be5842e7c883b74c5d9d0d57f98901cb8c656a8

Observation a2ab13f3-d9ea-46ec-836b-532c87a6817a · outbound

This paper cites Conversational contextual bandit: Algorithm and application.

Fusing Reward and Dueling Feedback in Stochastic Bandits Conversational contextual bandit: Algorithm and application

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:26:28.109362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T11:26:28.020817Z digest=sha256:e5986d4ee79a4a16b93b30d6179f392417e867d295efad3071a943042f1c7ac5

Observation 65e54e4e-bbd2-48a7-b69f-fb172ddfef67 · outbound

This paper cites write newline.

Fusing Reward and Dueling Feedback in Stochastic Bandits write newline

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:26:28.023566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:26:28.023566Z digest=sha256:48bf939392105045c4917e137012fb84211edeadb01b7836d131554648495e80

Pith citing papers

Observation 1fbed955-ec04-4fc6-8faf-ca4e1ae5031d · inbound

Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback cites this paper.

Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback Fusing Reward and Dueling Feedback in Stochastic Bandits

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:13.487443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T11:34:19.258512Z digest=sha256:5681dcb807bd8ac87427987fdb696f82ad1b411bff8c7f716dc276a8525aa455