Pith. sign in

REVIEW 1 cited by

Lock in Feedback in Sequential Experiments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1502.00598 v3 pith:KD5MBZVA submitted 2015-02-02 cs.LG

classification cs.LG
keywords methodsequentialexperimentationexperimentsfunctionsunknownwhenaddress
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We often encounter situations in which an experimenter wants to find, by sequential experimentation, $x_{max} = \arg\max_{x} f(x)$, where $f(x)$ is a (possibly unknown) function of a well controllable variable $x$. Taking inspiration from physics and engineering, we have designed a new method to address this problem. In this paper, we first introduce the method in continuous time, and then present two algorithms for use in sequential experiments. Through a series of simulation studies, we show that the method is effective for finding maxima of unknown functions by experimentation, even when the maximum of the functions drifts or when the signal to noise ratio is low.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Offline Policy Evaluation for the Continuous-Armed Bandit Problem

    cs.LG 2019-08 conditional novelty 4.0 of 10

    A δ-window extension of Li et al.'s offline bandit evaluation lets logged actions near a policy's choice count, giving a biased but rank-preserving (at coarse level) way to compare continuous-armed bandit policies.

Pith tools