Pith. sign in

REVIEW 3 cited by

Efficient Exploration for LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00396 v2 pith:E75O74GA submitted 2024-02-01 cs.LG cs.AIcs.CLstat.MEstat.ML

classification cs.LGcs.AIcs.CLstat.MEstat.ML
keywords explorationefficientqueriesagentfeedbackgeneratesuncertaintybenefit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present evidence of substantial benefit from efficient exploration in gathering human feedback to improve large language models. In our experiments, an agent sequentially generates queries while fitting a reward model to the feedback received. Our best-performing agent generates queries using double Thompson sampling, with uncertainty represented by an epistemic neural network. Our results demonstrate that efficient exploration enables high levels of performance with far fewer queries. Further, both uncertainty estimation and the choice of exploration scheme play critical roles.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. Personalized Recommendations via Active Utility-based Pairwise Sampling

    cs.IR 2025-08 conditional novelty 5.0 of 10

    A utility-based active sampling strategy for pairwise preference learning picks the questions that most improve expected recommendation quality, outperforming random and uncertainty-based baselines in two experiments.

  3. Large Language Model-Enhanced Multi-Armed Bandits

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Using an LLM as a reward predictor inside Thompson sampling and regression-oracle bandits outperforms LLM direct arm selection in the tested tasks.

Pith tools