Pith. sign in

REVIEW 2 cited by

Partially Observable Contextual Bandits with Linear Payoffs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11521 v1 pith:PVVYZF45 submitted 2024-09-17 cs.LG stat.ML

classification cs.LGstat.ML
keywords banditcontextualobservablebanditscontextsdecisionemkf-banditfiltering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The standard contextual bandit framework assumes fully observable and actionable contexts. In this work, we consider a new bandit setting with partially observable, correlated contexts and linear payoffs, motivated by the applications in finance where decision making is based on market information that typically displays temporal correlation and is not fully observed. We make the following contributions marrying ideas from statistical signal processing with bandits: (i) We propose an algorithmic pipeline named EMKF-Bandit, which integrates system identification, filtering, and classic contextual bandit algorithms into an iterative method alternating between latent parameter estimation and decision making. (ii) We analyze EMKF-Bandit when we select Thompson sampling as the bandit algorithm and show that it incurs a sub-linear regret under conditions on filtering. (iii) We conduct numerical simulations that demonstrate the benefits and practical applicability of the proposed pipeline.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents

    cs.LG 2025-05 reject novelty 6.0 of 10

    COBRA combines contextual bandits with a VCG-inspired leave-one-out detection mechanism so that truthful reporting becomes an approximate equilibrium while regret stays sub-linear.

  2. Scalable and Interpretable Contextual Bandits: A Literature Review and Retail Offer Prototype

    cs.LG 2025-05 reject novelty 3.0 of 10

    The paper reviews contextual bandit methods and sketches a category-level logistic-regression prototype for retail offers with LLM-generated member profiles, but provides no empirical validation.

Pith tools