Pith. sign in

REVIEW 4 cited by

Contextual Bandits for Unbounded Context Distributions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.09655 v2 pith:W6GFA554 submitted 2024-08-19 cs.LG stat.ML

classification cs.LGstat.ML
keywords methodalphabetacontextualregretunboundedbanditbound
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Nonparametric contextual bandit is an important model of sequential decision making problems. Under $\alpha$-Tsybakov margin condition, existing research has established a regret bound of $\tilde{O}\left(T^{1-\frac{\alpha+1}{d+2}}\right)$ for bounded supports. However, the optimal regret with unbounded contexts has not been analyzed. The challenge of solving contextual bandit problems with unbounded support is to achieve both exploration-exploitation tradeoff and bias-variance tradeoff simultaneously. In this paper, we solve the nonparametric contextual bandit problem with unbounded contexts. We propose two nearest neighbor methods combined with UCB exploration. The first method uses a fixed $k$. Our analysis shows that this method achieves minimax optimal regret under a weak margin condition and relatively light-tailed context distributions. The second method uses adaptive $k$. By a proper data-driven selection of $k$, this method achieves an expected regret of $\tilde{O}\left(T^{1-\frac{(\alpha+1)\beta}{\alpha+(d+2)\beta}}+T^{1-\beta}\right)$, in which $\beta$ is a parameter describing the tail strength. This bound matches the minimax lower bound up to logarithm factors, indicating that the second method is approximately optimal.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Research on E-Commerce Long-Tail Product Recommendation Mechanism Based on Large-Scale Language Models

    cs.IR 2025-05 reject novelty 3.0 of 10

    A hybrid LLM embedding plus attention plus score-fusion method is claimed to improve long-tail e-commerce recommendation recall and coverage.

  2. Deep Learning Model Acceleration and Optimization Strategies for Real-Time Recommendation Systems

    cs.IR 2025-06 reject novelty 2.0 of 10

    A standard combination of model compression and serving optimization gives 2.4x throughput on a GPU benchmark, but the headline claims of <30% latency and preserved accuracy are not supported by the paper's own data.

  3. Research on Personalized Financial Product Recommendation by Integrating Large Language Models and Graph Neural Networks

    cs.IR 2025-06 reject novelty 2.0 of 10

    A hybrid LLM-plus-GNN recommender is claimed to beat collaborative filtering, LLM-only, and GNN-only baselines on financial product ranking, with NDCG@10 of 0.372.

  4. LLM-Driven E-Commerce Marketing Content Optimization: Balancing Creativity and Conversion

    cs.CL 2025-05 reject novelty 2.0 of 10

    An LLM copywriting pipeline combining fine-tuning, vector search, and weighted reranking reportedly lifts CTR by 12.5% and CVR by 8.3%, but the evidence is unverifiable and internally inconsistent.

Pith tools