Pith. sign in

Paper Citation Record · LEDGER

Active Human Feedback Collection via Neural Contextual Dueling Bandits

As of 17 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2504.12016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12016 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:41:31.060383Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T01:06:19.756032Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T01:10:51.536288Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy45
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b92bfea-da5a-4e4e-b7b9-1be62bf2b948 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Improved algorithms for linear stochastic bandits

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.845727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.834866Z digest=sha256:0040496584eb25c98b853a78cb47c2c48e66a54b5b2fa05184ebf274c24bae69

Observation a9645d9c-3a57-4d78-b9b1-9f28b3ac4a87 · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Thompson sampling for contextual bandits with linear payoffs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.835873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.840112Z digest=sha256:9f209f5d389c62da165c2033fbac8e94390ac12fde20f6b1036d2cfec3179a86

Observation adcd636b-0d17-4df8-9b4a-6de545114d1b · outbound

This paper cites Reducing dueling bandits to cardinal bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Reducing dueling bandits to cardinal bandits

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.711781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.844259Z digest=sha256:ec6a8812fba5b2e55bba3ff3558fde0d861d79acbd36e95458fd3a78163bf45f

Observation d78e08d1-a64f-4d41-b23a-ee344c6c0475 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Finite-time analysis of the multiarmed bandit problem

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.701442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.847884Z digest=sha256:e403e9e7ee2969931bacbc194d3ceb00fefb446e511dfba23d595b8ed26df68c

Observation 7ec3452d-8f7c-472f-b818-d7c9dff37954 · outbound

This paper cites Bandits for Online Calibration: An Application to Content Moderation on Social Media Platforms.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Bandits for Online Calibration: An Application to Content Moderation on Social Media Platforms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.852505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.852505Z digest=sha256:b2ff6f7fcbab948f403428ce94127fe126465352ea162421e27eb43859f96e3e

Observation f96c0e42-fb07-4b39-93de-fadfb47162ba · outbound

This paper cites Neural Logistic Bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Neural Logistic Bandits

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.857457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.857457Z digest=sha256:8fdea94defe5700bc13ca9afd851fe63ba2aca5712e904fc2934f3c6f5a3cffd

Observation 52a4c249-e878-47d1-b9fc-f3addf71d6fa · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.862004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.862004Z digest=sha256:9b680d492bfb74c1f7fb7f15b523789cd2670febbd546dceb28f5383bdd5f9c2

Observation 9a43def9-9f82-4667-8dc9-267e16499be9 · outbound

This paper cites Ee-net: Exploitation-exploration neural networks in contextual bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Ee-net: Exploitation-exploration neural networks in contextual bandits

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.691294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.865789Z digest=sha256:a83ca74d8cf2a48a2b6537567caa46e824ba0394edd33a663606b0ba92b76fcf

Observation 7312d001-512f-43a1-b496-71e90252b08a · outbound

This paper cites Preference-based online learning with dueling bandits: A survey.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Preference-based online learning with dueling bandits: A survey

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.681049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.869175Z digest=sha256:2238560e7de940b4078e600083d66064d85bc06f423030b5141e04ae169846f5

Observation 8099a301-597b-4de6-8a3d-035f5a9d3650 · outbound

This paper cites Stochastic contextual dueling bandits under linear stochastic transitivity models.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Stochastic contextual dueling bandits under linear stochastic transitivity models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.670033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.872735Z digest=sha256:29c643045bdec06d3f8ddae9de580f15517d632f63b8423d1795c689b4c374fe

Observation b56997e0-b72f-45c1-bac9-b95c7a9464d7 · outbound

This paper cites An empirical evaluation of thompson sampling.

Active Human Feedback Collection via Neural Contextual Dueling Bandits An empirical evaluation of thompson sampling

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.658332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.877076Z digest=sha256:c5d9dacffe9c563407d52e9351a1f972f8f0fa21870ae64261bf85f111c90131

Observation 20c5dc23-7042-44d5-a059-4f7f8189e590 · outbound

This paper cites RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs.

Active Human Feedback Collection via Neural Contextual Dueling Bandits RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.880791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.880791Z digest=sha256:68cf7bcf34d7be8528f405b12e8d89d12b90adbfde9271eaeaa587cf3031738f

Observation bdb3c793-696d-4793-ba92-dbda4c0d9766 · outbound

This paper cites On kernelized multi-armed bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits On kernelized multi-armed bandits

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.646769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.884656Z digest=sha256:d7c23c994177032811d26618d679963c57231a83236cadb492f54543f535e27c

Observation b320fcd6-968d-4a54-8844-aad54f200d00 · outbound

This paper cites Federated neural bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Federated neural bandits

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.635328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.888391Z digest=sha256:2f2ce55dd2831d1c3d2d329a09cc1a006db8908ee9bec2d728fb1b14f90c6f55

Observation 8a06a7be-26ae-4b23-9a81-9672ddcaf82f · outbound

This paper cites Active Preference Optimization for Sample Efficient RLHF.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Active Preference Optimization for Sample Efficient RLHF

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.891792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.891792Z digest=sha256:e24ad21ce071b892b4644d7e887f481c93a230997768e9e1087f0196dc7ac4ce

Observation 38a00c37-13f7-452f-b3d5-8a2f868e4fdf · outbound

This paper cites Contextual bandits with online neural regression.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Contextual bandits with online neural regression

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.625397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.895545Z digest=sha256:beb39cddd68c4ef4170b28cfe36aab8351cff75ff69ef2c0f92a8e1195da6d19

Observation 29fb0a0f-9457-4bd4-9373-f539774e14a8 · outbound

This paper cites Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.899860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.899860Z digest=sha256:4e7803c7c99a2018a7318759d548e33023f0ab1d1eaf305852179ac57c3717e2

Observation 4b2709d2-84ff-4e9d-ac16-bd9157d54725 · outbound

This paper cites A relative exponential weighing algorithm for adversarial utility-based dueling bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits A relative exponential weighing algorithm for adversarial utility-based dueling bandits

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.615054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.903866Z digest=sha256:48c41c66c25b786a20b2dbb3bf0117ffc2af21d3549b4b3b21f18c8e9748d43e

Observation 8d2ac834-21f7-442c-ab0c-881fd6e2513a · outbound

This paper cites Mm algorithms for generalized bradley-terry models.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Mm algorithms for generalized bradley-terry models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.603611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.907383Z digest=sha256:fee0d283605c8476b5b5a179346446079a13c4ab46d4f0df7c0aaf1d14e04bed

Observation 1e3b336f-aaf8-4cef-906a-10f91b806d32 · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Neural tangent kernel: Convergence and generalization in neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.591942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.911587Z digest=sha256:5551bdb5a402ae0e5e1abe41563a4f22881c1bd53ffcba0dbc2eacc02b11da74

Observation 19c40104-86f5-48a9-a5f5-d6374f1e5341 · outbound

This paper cites Reinforcement Learning from Human Feedback with Active Queries.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Reinforcement Learning from Human Feedback with Active Queries

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.916115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.916115Z digest=sha256:8af5abbf48ee773fd2a7b85403b4fc38e8f625e906ed7ee8ce0b4b758414a2b5

Observation 264f5d06-b2b0-4837-af01-5ff28cccc185 · outbound

This paper cites A fast bandit algorithm for recommendation to users with heterogenous tastes.

Active Human Feedback Collection via Neural Contextual Dueling Bandits A fast bandit algorithm for recommendation to users with heterogenous tastes

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.579754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.919857Z digest=sha256:937976f6cd1a6b8cc9640f672bcd4e718e9dc513997e39cad5a8e8446fe0d8ef

Observation 8fe26830-3893-4cc3-b7cf-ae0509a9532e · outbound

This paper cites Regret lower bound and optimal algorithm in dueling bandit problem.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Regret lower bound and optimal algorithm in dueling bandit problem

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.568463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.923325Z digest=sha256:1cf518fe33d245b0f6e8effbfd2fe309919d15d6ce3976506cc6f23b83a465e0

Observation d4c74a89-3c93-4501-8f4b-ea2dec209c7f · outbound

This paper cites Asymptotically efficient adaptive allocation rules.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Asymptotically efficient adaptive allocation rules

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.556806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.927348Z digest=sha256:199f3f96f1546daa17f1004474babc7dfd177d801116dabd7df2b0333c73ba5e

Observation 090329ac-a71a-4489-b86b-9507d7079227 · outbound

This paper cites Bandit Algorithms.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Bandit Algorithms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.545036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.931922Z digest=sha256:7ba7c23cc8c1e57c5527ae93c232123f7166dc22cb92d75897b6659a7be7ee44

Observation 2140c30b-8451-4a54-a330-624278a32a88 · outbound

This paper cites Provably optimal algorithms for generalized linear contextual bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Provably optimal algorithms for generalized linear contextual bandits

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.534440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.935762Z digest=sha256:d9f8ab6e3c904d37c139e90a49d79dc133a2f3be6b7b9cd7b58f99d2e8125842

Observation 3c157c77-3234-4dfd-9690-f68fc4b83c0b · outbound

This paper cites Feel-Good Thompson Sampling for Contextual Dueling Bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Feel-Good Thompson Sampling for Contextual Dueling Bandits

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.940328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.940328Z digest=sha256:3953ee2348ebcad6c888d4c3d3584d8856b19e974096c02c29453fd89a917599

Observation 7bd464a2-91f2-436e-b4a0-a7980a855f51 · outbound

This paper cites Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.943994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.943994Z digest=sha256:41077a250731d7e38a49098e101e53479b952b26fffb86043812d044ea9ef55f

Observation 8300f739-f34b-460a-9c24-a4023fa59254 · outbound

This paper cites Prompt Optimization with Human Feedback.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Prompt Optimization with Human Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.947907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.947907Z digest=sha256:909f3ba766c95da3428993345fcb5a6b57424f16ffcfc237f168e4e4f4fd57a4

Observation 278c2e85-d364-49a7-a783-e2fbfa2d9904 · outbound

This paper cites Individual choice behavior: A theoretical analysis.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Individual choice behavior: A theoretical analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.951747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.951747Z digest=sha256:00bf895a1d41c66b56d70ee94b6aa54ef8c8d2638fa2b1faef2482e0cacccd7a

Observation 0ebf2c5f-4315-4abd-bcad-adf4dd5319ea · outbound

This paper cites Sample Efficient Preference Alignment in LLMs via Active Exploration.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Sample Efficient Preference Alignment in LLMs via Active Exploration

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.955696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.955696Z digest=sha256:4cf39aa737b1883b3e84de90f148d03820aa598be11ced8c613d84d9be773d07

Observation 2069cad4-d689-406c-afde-9eb895806c61 · outbound

This paper cites Teaching language models to support answers with verified quotes.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Teaching language models to support answers with verified quotes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:30.959827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:30.959827Z digest=sha256:97471311ea393517d6200c6a3a12061cf503f3ff5003246951e695c7af4a3d9f

Observation 82d87dc1-be05-49ef-8102-dff7346f7045 · outbound

This paper cites Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.517327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.964513Z digest=sha256:fec359ac97ab19640d8b3bbf33de897fa52d6f585ca9d6dc7737a65f4fe69e84

Observation 9c5d88ae-ffac-4c96-9f95-a37c03de196d · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Optimal algorithms for stochastic contextual preference bandits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.506725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.968362Z digest=sha256:cd7bbbdf9fc1066278e1730c7665f90662e359c142ca316ba07c3e7f038ce752

Observation 93427b48-f9a3-4c15-8013-0f3eeca7e165 · outbound

This paper cites Battle of bandits.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Battle of bandits

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.496166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.971875Z digest=sha256:7ad47d7467cd9c14d8636ef300557adbaab994e2872bb3ac72012b609059ba2c

Observation 39c6f195-ae93-4208-8fe3-5b4fe1c3c5e3 · outbound

This paper cites Active ranking with subset-wise preferences.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Active ranking with subset-wise preferences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.485691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.975493Z digest=sha256:93e58df51029c18a893577acecd2e5e82b1905982ab2d7c4e7e7a39c6992bcb4

Observation 98d39aea-4964-4a15-bf8d-606e9f43d0f1 · outbound

This paper cites Pac battling bandits in the plackett-luce model.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Pac battling bandits in the plackett-luce model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.474812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.978837Z digest=sha256:4f06df55f88fa78f0992cebd8abcc9058099c562ea2de7ff7ec8509a0dc512c4

Observation 9d5c356e-0aaa-44ce-a094-46d3311f01fe · outbound

This paper cites Efficient and optimal algorithms for contextual dueling bandits under realizability.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Efficient and optimal algorithms for contextual dueling bandits under realizability

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.464377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.982868Z digest=sha256:63bb7d13bf36bb3d5ea4aebbe60c2d9b66579010bae808d3717258b7631210ca

Observation eb47f4f3-89ee-4281-8013-9840548d9a26 · outbound

This paper cites Computing parametric ranking models via rank-breaking.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Computing parametric ranking models via rank-breaking

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.453961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.986131Z digest=sha256:405f2495eb9c35f8aede387adf9ffff36b7172fd6856458225967f47baad1af9

Observation 664ec347-ec3f-4de6-b9fe-ea9d55c6f0a8 · outbound

This paper cites Gaussian process optimization in the bandit setting: No regret and experimental design.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Gaussian process optimization in the bandit setting: No regret and experimental design

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.442321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.989299Z digest=sha256:a04d4de737c70f51d9d9344b18cd83e764038b444d12b0114757c9d9deaca0f5

Observation ee242278-69c0-46d2-8d3f-851be3763a74 · outbound

This paper cites On the likelihood that one unknown probability exceeds another in view of the evidence of two samples.

Active Human Feedback Collection via Neural Contextual Dueling Bandits On the likelihood that one unknown probability exceeds another in view of the evidence of two samples

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.429551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.993481Z digest=sha256:e1cc878c7367fe8bd2bb8d2d6edd393831112ef6b4bbf6a01ec2cfb1c435f7a3

Observation e94b8650-1298-4d6a-8a04-65ae2a7c7244 · outbound

This paper cites User-friendly tail bounds for sums of random matrices.

Active Human Feedback Collection via Neural Contextual Dueling Bandits User-friendly tail bounds for sums of random matrices

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.417330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:30.998088Z digest=sha256:95c1f24b0b786cfa08c8d05871988d001cf4ca9a98312313bc589a3dcc752cec

Observation 8e9f64dc-d676-456e-a592-dd5078da59d3 · outbound

This paper cites Online algorithm for unsupervised sensor selection.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Online algorithm for unsupervised sensor selection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.405276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.001315Z digest=sha256:b2f7814da8268e9e46a9d41e37c1873d27d57587866f026c78c4815a334df789

Observation 5ce20994-2cb4-4923-b847-97ed38d9bd5f · outbound

This paper cites Thompson sampling for unsupervised sequential selection.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Thompson sampling for unsupervised sequential selection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.394635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.004742Z digest=sha256:6d8319f800459ab28f3e99fc5cf177856a217fef212099c66baf986617af8d37

Observation 79504d2e-c551-41c1-b322-9d305db06b91 · outbound

This paper cites Online algorithm for unsupervised sequential selection with contextual information.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Online algorithm for unsupervised sequential selection with contextual information

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.383866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.007879Z digest=sha256:230557dff01f18722ce3d49398b577b04e53f7704da29a3c273e6723345f0d4d

Observation 4f28bc8c-24cc-47ec-92fa-aa8d5637d839 · outbound

This paper cites Neural dueling bandits: Preference-based optimization with human feedback.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Neural dueling bandits: Preference-based optimization with human feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:41:31.011444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:41:31.011444Z digest=sha256:5a34105f956f98341180def37d426326e8110f9f757f19dcd32d87f5ccc40b8c

Observation 6e706e87-c5bf-43ac-8bcb-806d9ee62a2a · outbound

This paper cites Gaussian Processes for Machine Learning.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Gaussian Processes for Machine Learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.365507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.015178Z digest=sha256:045f83a7f3b7cfda24a423a0edb4d1f2193adfb7e420e275f108d2c84039d157

Observation d4304843-3dec-4557-b4c3-5e41b9f2fbb2 · outbound

This paper cites Personalized news recommendation: Methods and challenges.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Personalized news recommendation: Methods and challenges

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.355016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.018904Z digest=sha256:eb2b8acd89777b4f79c4c81bceed8a2b0fa039c1773135bda60af0cc211bdd4a

Observation d5ee8ddd-2d55-423e-9b24-8198a054966f · outbound

This paper cites Neural contextual bandits with deep representation and shallow exploration.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Neural contextual bandits with deep representation and shallow exploration

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.344204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.022322Z digest=sha256:249e65c8e78cc65b9215d8717e431f6b340c447bd161492e1766c33077e42e28

Observation b8dc5c22-932e-4233-8114-eb2ba34e2e71 · outbound

This paper cites Conversational dueling bandits in generalized linear models.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Conversational dueling bandits in generalized linear models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.332415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.026274Z digest=sha256:fb2d602b8e604848b2351e76235ad61218c67b5f0a18713ce20e826502a57f40

Observation 3983f234-323c-40a4-a63d-42b824989da0 · outbound

This paper cites Interactively optimizing information retrieval systems as a dueling bandits problem.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Interactively optimizing information retrieval systems as a dueling bandits problem

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.319415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.029949Z digest=sha256:ce7ffd1ff5aa051f35cf1bfddb3d8db5ac78f499171ea04974d0d8085db7b717

Observation 2b0c2085-7d11-4cb6-9588-487fbbe90474 · outbound

This paper cites Beat the mean bandit.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Beat the mean bandit

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.307710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.033340Z digest=sha256:29a97b7f9d6e4866fb3a83aa195c932be11257410d415c4c4d62bdbd9a6336aa

Observation 211bb2f3-2520-4da6-927e-ed0338cea2c0 · outbound

This paper cites The k-armed dueling bandits problem.

Active Human Feedback Collection via Neural Contextual Dueling Bandits The k-armed dueling bandits problem

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.294853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.037869Z digest=sha256:a230035ed1d85143afe919ce1c83c74ed3afe95cfe3f548563ba9ee51c9940ef

Observation 74ecfffa-3239-48b9-b6e9-e4fae2f6a09f · outbound

This paper cites Neural Thompson sampling.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Neural Thompson sampling

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.282920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.041757Z digest=sha256:528b522ba1fc36f72766f77e035aa00aa227a7eca674cb8b06c83e85944c02db

Observation 2e9a3248-f4e1-45de-aabe-3b9ab50b412c · outbound

This paper cites Prompt learning for news recommendation.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Prompt learning for news recommendation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.271839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.045186Z digest=sha256:eab0754a8d8ff271579a6301e0c726d0b73e29112ac17b787b9471fc089f78de

Observation 0390bddf-6d4e-4f22-b23f-ea587fc15d69 · outbound

This paper cites Neural contextual bandits with UCB -based exploration.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Neural contextual bandits with UCB -based exploration

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.260521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.049962Z digest=sha256:79a81dfb131a3f974ee76bf0102fd143661d0037f0fabfdc87dd3500ad4ac5f1

Observation 81938446-e3cc-4408-988e-65d45bf51ea6 · outbound

This paper cites Principled reinforcement learning with human feedback from pairwise or k-wise comparisons.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Principled reinforcement learning with human feedback from pairwise or k-wise comparisons

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.249600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.053242Z digest=sha256:45061bf039310106a786751ad2796f2b3df71ac3eb47785512ab59036ab46f39

Observation 56ec91cb-4348-4cc5-ad2f-539cd85d4c6c · outbound

This paper cites Relative upper confidence bound for the k-armed dueling bandit problem.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Relative upper confidence bound for the k-armed dueling bandit problem

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.237658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.056756Z digest=sha256:569a66c283953db63dd25ee2882ce491252c1fa17255100abe881fd72e20a73f

Observation 28b099bf-4a80-4cf7-a92f-413e2588e34f · outbound

This paper cites Relative confidence sampling for efficient on-line ranker evaluation.

Active Human Feedback Collection via Neural Contextual Dueling Bandits Relative confidence sampling for efficient on-line ranker evaluation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:41:31.224928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:41:31.060383Z digest=sha256:770a7c1bb6d8c7ff80247bbe83dde279e76a06d513d7100c281319c256646af9

Pith citing papers

Observation 1c677bae-ddef-450e-865d-e26463d7a187 · inbound

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment cites this paper.

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment Active Human Feedback Collection via Neural Contextual Dueling Bandits

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:10:51.539069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T01:06:19.756032Z digest=sha256:5aaf89efa47c0a1ee41fd8de510bbaf1957dbae8aa18d4ef9c4d3905579f7bc5