Pith. sign in

Paper Citation Record · LEDGER

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.23673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23673 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:45.634287Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbb2eea2-eefa-4813-a6ce-abe16523ff2b · outbound

This paper cites write newline.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:39.878545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:39.878545Z digest=sha256:88a45c9c41704e098543841097a85ad66e3dc1ba0aaf9727d4c25a527a2041f2

Observation f25d9bbb-539f-4ffa-a943-5380cdd1e2ee · outbound

This paper cites Online learning for linearly parametrized control problems.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Online learning for linearly parametrized control problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:57.055986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:39.976257Z digest=sha256:67e3c4f5dfe9eb9a2752c165c95fd7de1379e84d9cbb0e30192837b94551fd49

Observation f63f9162-3596-4bc2-b4e7-b14c5fb6ed7d · outbound

This paper cites Reducing dueling bandits to cardinal bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Reducing dueling bandits to cardinal bandits

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.865663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.043920Z digest=sha256:44bae7e3f2ded93131972aba692ca22fd29cba7bf703b36b329cb8e204b888f8

Observation c916829b-9808-4ebb-a0ce-b4bda5cf7a20 · outbound

This paper cites S., Hu, W., Li, Z., Salakhutdinov, R.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds S., Hu, W., Li, Z., Salakhutdinov, R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.138320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.138320Z digest=sha256:73f0364283c2df90a9e355afaae8fc19acc1861375724948d06dc0993cf3e18e

Observation 661e7698-f079-457d-91a0-dcbd759a411c · outbound

This paper cites R., Daulton, S., Letham, B., Wilson, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Daulton, S., Letham, B., Wilson, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.750917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.241854Z digest=sha256:0ddd072560cabace017aa415b44891b7e99c005bda09c49608ef873a1460c6ae

Observation 3c8de591-d8e2-458b-9040-58d0cb36bd85 · outbound

This paper cites Preference-based online learning with dueling bandits: A survey.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Preference-based online learning with dueling bandits: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.449269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.361927Z digest=sha256:4f56b4aaa4524f4582cfc8d72a7ff09760ec51e4df284654ebca18eef2245efb

Observation 3b08b2bb-24a7-42e8-af6a-df979b24cc66 · outbound

This paper cites Stochastic contextual dueling bandits under linear stochastic transitivity models.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Stochastic contextual dueling bandits under linear stochastic transitivity models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:56.187735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.442785Z digest=sha256:c2e0332c5d7d3e0aaef529aff42342a3956dd332bd5511a360aa6f06c7f3885f

Observation c5f9d829-ad58-45fd-b147-a251fe5e750a · outbound

This paper cites Mat \'e rn gaussian processes on riemannian manifolds.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Mat \'e rn gaussian processes on riemannian manifolds

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.922717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.521123Z digest=sha256:b8fe1382e2c858f397515dfbfd9ce4c86ffc95ca979567ea074611b90d7b3340

Observation 1cbd7aa1-5dfe-4a1d-9f66-859d3c207769 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.628781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.628781Z digest=sha256:8d043b2ef09dba5a510cdb66cc3d5747e20593a74ffb0223c2eef73af4e0721a

Observation d1b04aac-25ff-468c-a4f1-279ce8a31e28 · outbound

This paper cites A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.734979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.734979Z digest=sha256:3b77a7a0413642567128369be75fd9fb3ae8d6c531efa40c675cdd782d9fdda0

Observation eb9496c5-f672-4075-a016-e37658ba9289 · outbound

This paper cites Instructzero: Efficient instruction optimization for black-box large language models.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Instructzero: Efficient instruction optimization for black-box large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.729489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.835746Z digest=sha256:1fae537e3498fc596aac1ba54a94000b743536b95cd8e5a84163db693abc6b1f

Observation 72f068b0-e951-4dec-aabc-ef40b7e314d5 · outbound

This paper cites Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Human-in-the-loop: Provably efficient preference-based reinforcement learning with general function approximation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:40.901099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:40.901099Z digest=sha256:5a0a1a0cc1f9557252e336cb8cc2659d2ba074e7585673562c3f8191ddacb6e4

Observation cd4f5627-73e6-42a9-9a40-4be5c0d094e4 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:55.354090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:40.994901Z digest=sha256:c18f36b629f907d7ec4b7f095aaed271946350e2b47c2c37ad1aa20651b4dd93

Observation 3a2067e6-4446-4226-8812-ca1d198dd0cc · outbound

This paper cites and Steinwart, I.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Steinwart, I

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:55.009181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.069695Z digest=sha256:6ef44840b6e409f9d25f5ba614bb0bcb29560468baff220c8c7c0684a6ea7c5c

Observation bf1ae0a8-8f2e-4f09-9963-fb340e6a3851 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:54.743049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.168084Z digest=sha256:e8733b2c9fbe4021676a17436e8ef494278a606428a91de0faa19f4d92d01aeb

Observation 6dd47c75-eda3-44a6-93c1-9020fbd86d56 · outbound

This paper cites E., Slivkins, A., and Zoghi, M.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds E., Slivkins, A., and Zoghi, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:54.433457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.269844Z digest=sha256:51b330df2fc936e9864e20d4a446a4a7054413e37be7d2a699d3a38c11b83a6c

Observation 279f94d4-b194-4864-b290-d74173a76ca7 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:54.070394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.351047Z digest=sha256:97c2cf2f962c2c601dd0b6ebbceb166be148c849ee26bd7310d09d36dae0861d

Observation 90cc3a77-2405-4f42-bf4c-0d4f49d85df7 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Improved optimistic algorithms for logistic bandits

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:53.772960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.453532Z digest=sha256:aaa1652d130f1f5cf056c3a75cc2b440f55b9e2dd3db9d963495cd0ce84f6202

Observation 3fd1dca9-b811-4db7-8ba6-edaa73054cf9 · outbound

This paper cites A Tutorial on Bayesian Optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A Tutorial on Bayesian Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:41.544682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:41.544682Z digest=sha256:fd07e74b032829acc1e592054767a07380d033c74b5699a0ac96d13a74f36879

Observation 18fce1e8-13d2-40ae-9586-1d9ec0b24e46 · outbound

This paper cites R., Pleiss, G., Bindel, D., Weinberger, K.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Pleiss, G., Bindel, D., Weinberger, K

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:53.480243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.638418Z digest=sha256:f9eb72704622db0e24ae954a1c982932dbd6362fc766e8103bba610436f3dbf5

Observation b5164a3c-46a1-4518-9082-846da787c71c · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:53.193134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.704728Z digest=sha256:e526d870a123edc85c94b6b67cba8216841295c716d5efa61a266e091c53c602

Observation 1cc4fddb-d77a-422f-bbb7-62bba9674992 · outbound

This paper cites L., and Thomaz, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds L., and Thomaz, A

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.899224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.773886Z digest=sha256:12161cabfe954814add3c3e2e6f374cd26de3be76f51847aded17455c2358522

Observation f82de9f4-ae72-47bb-a88c-9b04475e04bf · outbound

This paper cites and Yang, X.-S.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Yang, X.-S

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.600042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:41.930302Z digest=sha256:4b0def897e726bb3a180e7befedca7a0473a4541cfe3324973177df7a54424d6

Observation dad41594-2365-4926-9a1e-5a488c4bb929 · outbound

This paper cites R., Schonlau, M., and Welch, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds R., Schonlau, M., and Welch, W

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:42.054479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:42.054479Z digest=sha256:645f6010fb02aa409c924f0fc3272f77b40d8c03b69951d6f7b9224046e3f76f

Observation 807202dd-e4de-4525-b1ca-f3f9eeb24320 · outbound

This paper cites Feel-good thompson sampling for contextual dueling bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Feel-good thompson sampling for contextual dueling bandits

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.335804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.145030Z digest=sha256:2c65e363162c55f5793f39322a0c78e43db87ed8a2a37665a5befcbbf636d51d

Observation 70eed8b9-bb8e-4b6d-b810-cfc21206ee71 · outbound

This paper cites and Scarlett, J.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Scarlett, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:52.109171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.274206Z digest=sha256:d00f3967f09cece1b0f2ad6cde6a0edc77cda06f85f5b36dc3ffbf78baa11479

Observation 6fa7c559-fe37-473b-916b-3147b64858b7 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:51.870471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.368045Z digest=sha256:34bbd7144cdca3dfc8fed418addda057d1ea075691b14c46aa3513e82474df6c

Observation c52c829c-9edc-4326-9bd9-17ec706ccd74 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:51.603727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.497274Z digest=sha256:6153d8e8d568a890bd0569abb2c1a24169b9d5f84ff0fa1ba785f6fa9fb194a1

Observation 3e05f0f4-440a-4c96-815a-b223275df31c · outbound

This paper cites Sample Efficient Preference Alignment in LLMs via Active Exploration.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Sample Efficient Preference Alignment in LLMs via Active Exploration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:42.595875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:42.595875Z digest=sha256:542b16b07a30d332fb37fc8e5621e33119351eb4c0c7213a739c29d714c66e80

Observation 05359119-0546-450b-85ca-779f84c2d013 · outbound

This paper cites Kernelized offline contextual dueling bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Kernelized offline contextual dueling bandits

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:51.302536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.722443Z digest=sha256:e28924172c5e148f8aee9f639a73e7fdf28311ebb56db038d69fd3cfa5e4c188

Observation 5a862fc4-f9e7-4c39-a3ff-b6c87c6ebf2d · outbound

This paper cites Functions of positive and negative type, and their connection with the theory of integral equations.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Functions of positive and negative type, and their connection with the theory of integral equations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:51.069454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.844443Z digest=sha256:b3af3e219a2f763c0ca244b92088e7269701be0855dc707e382aac0f5966183e

Observation 1ce9876e-4bd5-4bb8-86be-301ac54aa385 · outbound

This paper cites Projective preferential bayesian optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Projective preferential bayesian optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.884143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:42.936733Z digest=sha256:9720585cbfee5c1062136dc2ea9118380fcd13a899c6a98463af82b5bdd038a1

Observation 849b6473-3f9d-4ee8-883c-119846cf8701 · outbound

This paper cites Dueling posterior sampling for preference-based reinforcement learning.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Dueling posterior sampling for preference-based reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.040733Z digest=sha256:761c28c3c851ced8f0e960c64afb2860d55800bdcb2d4176d79cf5ad701cb659

Observation 8646b46f-3fbe-4102-89a5-ea926abb3064 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.162585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.162585Z digest=sha256:c3a20f0b885e8956f8e0e4b21a7cb683c8ed4f8ad988bbb578979cff1b18dc7d

Observation 4f7d64a3-69ac-49b3-a6b5-ab2d59de08c1 · outbound

This paper cites Bandits with preference feedback: A stackelberg game perspective.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Bandits with preference feedback: A stackelberg game perspective

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.681901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.268258Z digest=sha256:ce62d78c8e3b0156a0895a12c9937082c9539c98489bf0399651b68508d5d686

Observation caa9fe46-d605-42c5-bc2a-bcebf3c6eac5 · outbound

This paper cites Scikit-learn: Machine learning in P ython.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Scikit-learn: Machine learning in P ython

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:43.415191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:43.415191Z digest=sha256:95a557de3a9578cc2050bf16ec617f0a19b39045a5e09307115ee73ff0a242f8

Observation 166d52dd-32ae-4492-b81d-5df8ea1d1138 · outbound

This paper cites Optimal algorithms for stochastic contextual preference bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Optimal algorithms for stochastic contextual preference bandits

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.504449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.528257Z digest=sha256:d031feb406a47a39efec397f89938bf174cb5d3da9ced522ac447ae5d87e7ddd

Observation 7d52213e-1a68-489b-9e55-ee72fe4bb353 · outbound

This paper cites and Krishnamurthy, A.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Krishnamurthy, A

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.376160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.638184Z digest=sha256:b9c18818c2c6eaf9332479f9da3964b8c63c3d69b80243045188874177111dae

Observation 889bb251-af6e-403a-9647-950bd1ee4e92 · outbound

This paper cites Dueling rl: Reinforcement learning with trajectory preferences.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Dueling rl: Reinforcement learning with trajectory preferences

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.226512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.712472Z digest=sha256:a1cf87bfdc21962339d3b5b278343b737e93e0ecdfeae9c193e3a000ca5c993d

Observation 1c144d77-19f3-46d8-b247-76c1035bb965 · outbound

This paper cites A domain-shrinking based bayesian optimization algorithm with order-optimal regret performance.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds A domain-shrinking based bayesian optimization algorithm with order-optimal regret performance

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:50.071549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.782700Z digest=sha256:bb11b1363f73843c2aa54f91cf2607df774295ed6647f2bfcf8709d17d2e4c52

Observation c2c9467f-1cca-4af2-bb75-07d3faab1062 · outbound

This paper cites Lower bounds on regret for noisy gaussian process bandit optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Lower bounds on regret for noisy gaussian process bandit optimization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.875455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.853547Z digest=sha256:e8881facbe68679b3b5462218fbfcfd202a1f8ada330ccb7ad82547a864efc2d

Observation b123e8fe-ebfc-4bcf-940c-41a2afa6489c · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:49.701942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:43.935161Z digest=sha256:cd75b9f3846c12cb5e5c9eefe40e88a7a6d878a604939ce42981bb4b548e67ef

Observation 3758a7e0-2e92-465b-bbf7-cdeb81572b1c · outbound

This paper cites P., and De Freitas, N.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds P., and De Freitas, N

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:44.036902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:44.036902Z digest=sha256:9234d725ebf93ba87bcb996d9c4b537087c880beb3394f4c338884a4a278a7b3

Observation 23b44e2a-b881-4f5f-981d-8fbcc8ea02c5 · outbound

This paper cites M., and Seeger, M.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds M., and Seeger, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.499940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.117889Z digest=sha256:1782484d6953ae364da37aad415719cef2aa9526a218f061093cd21bd99e8f1b

Observation e8d65b47-20a1-4c62-9db7-57bebf343634 · outbound

This paper cites Towards practical preferential bayesian optimization with skew gaussian processes.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Towards practical preferential bayesian optimization with skew gaussian processes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:49.311254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.202090Z digest=sha256:c5d6fe291460d7eb3b0747b9941e91ce23b1b5c02832d2b636350b423d1b552e

Observation a985fd11-9b1a-4b58-a2eb-372ec3ca4b69 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:49.159137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.270455Z digest=sha256:dd38d71c0e50cb6fb8f1a22d678d4d38de6ea8db7c4dc8c822d1570b68e40b11

Observation 08434648-ed7d-40f3-bf5d-6b6810b15521 · outbound

This paper cites Optimal order simple regret for gaussian process bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Optimal order simple regret for gaussian process bandits

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.991095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.350553Z digest=sha256:160637f37f67f58096853c3520e661757c579e36a95a282f4f04ddb7521a969a

Observation b9381b72-7c92-4d46-9a8c-8c9188bc71d4 · outbound

This paper cites On information gain and regret bounds in gaussian process bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds On information gain and regret bounds in gaussian process bandits

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.847219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.435530Z digest=sha256:64723d6014f027b6f5e9d9609b0ea644c97e0528f6c84b64040a2ed97ca2625a

Observation 4d154dbc-42cf-4bea-8214-ebb4695eacbe · outbound

This paper cites Finite-time analysis of kernelised contextual bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Finite-time analysis of kernelised contextual bandits

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:48.739454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.531719Z digest=sha256:1fd9f866c97f5130c70f1f8fdf8c107e078a0a0d0f6f26c32807650f6a59eb81

Observation b91b39b6-17ce-4ffc-a485-0509836099e9 · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:48.477948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.601543Z digest=sha256:c7368003c838456afb9184b1a961cbf290dc52f3c1bfd7abe5540bbd4c1a9652

Observation 4fd5c7cf-d33e-4877-9957-135aea590e2a · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds High-dimensional probability: An introduction with applications in data science, volume 47

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:44.680747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:50:44.680747Z digest=sha256:466839ed40a770a8358844428ee8a7866844c66ae01d23ec6dd30b94035500fd

Observation 25dd1c45-fc65-4544-b0e1-c48f8c1ee64d · outbound

This paper cites an unresolved cited work.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:48.210392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.789087Z digest=sha256:3db0252902d2217433403b32906d742a2008edc2d85f7860713a14c7a9e58c7b

Observation 6e71dfec-1ae2-40e0-934d-9d0c8eee30b4 · outbound

This paper cites and Sun, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Sun, W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.955298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.884469Z digest=sha256:b34e74bbc130f7dcbaff564df9061aae054fb0db66704bf9bc1b09917fd0b66b

Observation 92e305af-7512-4a3b-85d2-659758cb44ae · outbound

This paper cites Principled preferential bayesian optimization.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Principled preferential bayesian optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.700299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:44.981051Z digest=sha256:fc336e69f797ce17518295b8ae83dc6f603c21da41cb874bc701be6d0f29177a

Observation bb3d5c7e-ab48-4452-b719-d08bd5b063cb · outbound

This paper cites Zeroth order non-convex optimization with dueling-choice bandits.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Zeroth order non-convex optimization with dueling-choice bandits

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.449181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.103390Z digest=sha256:cc393d4b8b23e3d83399fef97b0c2e475069a1660217fb0c3d365dd7066ba25f

Observation f840f36d-5b10-4be3-8f71-ee7fb1ae6126 · outbound

This paper cites Preference-based reinforcement learning with finite-time guarantees.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Preference-based reinforcement learning with finite-time guarantees

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:47.176236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.192620Z digest=sha256:8696d504fb1b1abab04b11f3f3746cd04e2572a60735d605529d1cdf535d0238

Observation 9f09a41f-a9d6-47d5-91eb-6c09adb7e6d7 · outbound

This paper cites and Joachims, T.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds and Joachims, T

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.923255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.253187Z digest=sha256:6640e04803bf4f4510ea5070af84e7c7b4766016f2fc48198652a4387fc0e740

Observation f90a452d-7f1b-4882-ade5-a0cca8fe9794 · outbound

This paper cites The k-armed dueling bandits problem.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds The k-armed dueling bandits problem

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.669843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.318578Z digest=sha256:bf5a3f4e19df4b88a7c3f9eda8f20ca6e0b901459c49fa262f0e3dc46a9742f3

Observation 084eb858-f460-4be7-b7e6-405729726a10 · outbound

This paper cites D., and Sun, W.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds D., and Sun, W

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.427524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.418047Z digest=sha256:5a8a72414f0ca8b6ff083373accbbeda74344e0254489b62e7b9a80aa65dc12a

Observation 9d3dc523-5eee-4577-b38a-cb2a093b2b88 · outbound

This paper cites Relative upper confidence bound for the k-armed dueling bandit problem.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Relative upper confidence bound for the k-armed dueling bandit problem

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:46.179608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.531566Z digest=sha256:99ad2759b68c8edb60bf77da8a0571d86f79ad23c8ada89ca845a6a55c15aed6

Observation 243d5c50-0041-43b5-9901-9c4cb9c5cdb0 · outbound

This paper cites Mergerucb: A method for large-scale online ranker evaluation.

Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds Mergerucb: A method for large-scale online ranker evaluation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:45.950061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:50:45.634287Z digest=sha256:8626b32e22d09813f3fc58a1b9213647efdb7e6bbd49515e2fb1c5db24bd6602

Pith citing papers

No inbound Pith citation observations are available.