REVIEW 2 major objections 5 minor 79 references
Adaptive Bayes exactly tracks information over intrinsic time
T0 review · 2 major / 5 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read Bayesian and multiplicative-weights updates pay an exact information ledger: excess loss equals a round’s uncertainty cost plus a drop in distance to the comparator, measured on a pathwise clock called intrinsic time.
desk verdict Exact pathwise ledger for variable-temperature Bayes/Hedge is real and carefully bookkept; breadth claims and luckiness rest on stated premises, not on a broken core. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The one-step information balance: excess composite loss equals the centered mixability gap plus a scaled drop in KL divergence to the comparator. Composed two ways, it yields the retempered identity with drift and terminal free-energy terms, and the local identity with cumulative normalization and terminal mass.
What would settle it
On a fixed synthetic loss path, recompute the three terms of the retempered decomposition and check whether their sum equals composite-loss regret to machine precision for several comparators and schedules; a systematic residual larger than numerical noise would refute the identity.
Extended reading notes
Core claim
For any predictable positive learning-rate schedule and any comparator distribution, the composite-loss regret of a prior-retempered Bayes update equals temperature-change drift plus terminal comparator information plus the sum of learning-rate times per-round intrinsic-time increments; a parallel exact three-piece identity holds for the local pressure-target update. The per-round increment is the nonnegative finite-temperature cumulant of the played distribution, not a proxy variance.
Load-bearing premise
Each round’s exponential normalizer must be finite at the chosen temperature; for the fast expected-rate claims, a strong low-noise condition around the comparator must also hold, and the paper notes that condition is usually empty for diffuse mixtures.
Editorial extensions
If this is right
- Favorable stochastic or low-noise sequences appear as self-bounding intrinsic time inside the same pathwise identity, without separate algorithms or proofs.
- Side information, optimism, and compensators only change the residual sequence fed to the same Bayes update; improved regret is reduced unexplained information.
- Schedules can be designed from the revealed clock (square-root on cumulative intrinsic time, or pressure targets that hit a one-step free-energy level) rather than from a known horizon.
- The same exact ledger transfers unchanged to boosting margins, continuous-action online convex optimization, contextual bandits with estimated losses, and repeated-game regret matching.
- Classical first- and second-order regret bounds are successive relaxations of one exact cumulant term, so the order is an analyst choice, not a different algorithm.
Reading between the lines
- If the ledger is truly schedule- and geometry-agnostic once composite losses are fixed, many “new” adaptive experts algorithms may amount to different controllers on the same two update cells rather than new proof objects.
- The pressure-target view suggests treating one-step free-energy calibration as a thermostat: non-equilibrium exchange relations could give exact fluctuation identities for local updates.
- A practical diagnostic for deployed sequential reweighting (including preference and post-training pipelines) is to plot the three share terms over time; equal terminal regret can hide very different pay/drift/info mixes.
- Extending simultaneous single-copy quantile adaptation while keeping the exact cumulant under the played distribution would close the remaining gap between fixed-budget PAC-Bayes control and fully parameter-free scaling-time results.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops an exact pathwise information-accounting identity for Bayesian and multiplicative-weights updates. On each round, excess composite loss to any comparator ρ equals an immediate centered cumulant payment δ_t(c) (or η_t Q_t(c)) plus a KL transport term. Composing these one-step balances yields two exact three-piece cumulative decompositions: for the prior-retempered update, R_T^c(ρ) = D_T + B_T(ρ) + Σ_t η_t Q_t(c) (Theorem 2.10 / Eq. 18), with intrinsic time V_T(c) := Σ Q_t(c), temperature-change drift D_T, and terminal comparator information B_T(ρ); and a parallel identity for the local pressure-target recursion (Proposition 3.13 / Corollary 3.10). The same calculus is applied to side information, shifting/quantile comparators, stochastic luckiness, continuous OCO, boosting, bandits/feedback graphs, and repeated games. Empirical diagnostics in §7 and Appendix C report machine-precision residual checks of the prefix identities and envelope tightness.
Significance. If the identities hold as stated, the contribution is a genuine unifying ledger rather than another family of upper bounds: favorable regimes appear as self-bounding of realized intrinsic time, and many standard dichotomies (first- vs second-order, hard vs easy sequences, full vs partial feedback) become different readings of the same exact split. The manuscript supplies the main algebraic proofs in Appendix B, machine-precision residual checks of the prefix identities across (K,T) grids, and a clear 2×2 design space (RET/LOC × SQRT/PRESS). That combination of exact bookkeeping, schedule design from the same clock, and empirical verification of the identities is a real advance over variance-proxy analyses that introduce Q_t only after inequalities. The breadth of applications is secondary to the core ledger claim, which is load-bearing and carefully scoped to finite one-step log-normalizers.
major comments (2)
- §4.5, Condition (74) and Theorem 4.9 / Corollary 4.11: the comparator-centered low-noise condition is load-bearing for the constant expected-regret claims, yet the paper itself notes it is practically vacuous for most non-degenerate diffuse posteriors (forcing near-deterministic losses on the support of ρ). The fixed-rate and predictable-rate luckiness theorems therefore deliver meaningful constant rates primarily for point-mass comparators. The abstract and §1.5 still present “favorable stochastic or low-noise regimes appear as self-bounding properties of the realized intrinsic time” as a general selling point of the exact decompositions. Either restrict the fast-rate claims explicitly to point comparators (or highly degenerate environments) in the abstract/contributions, or supply a non-vacuous condition that covers diffuse ρ without collapsing to the point-mass case.
- §1.5 item 4 and §5–6: the claim that “the same calculus covers … continuous priors, boosting, online convex optimization, contextual bandits, and repeated games” is true at the level of formal transfer of the one-step balance, but several extensions are thin. Continuous-action OCO (Theorem 5.11 / Corollary 5.12) relies on an expensive density update and barycenter whose computational cost is acknowledged but not quantified; the bandit section (Theorems 6.1–6.6) correctly isolates martingale and bias terms yet does not report numerical residual checks of the estimated-loss identity comparable to the full-information checks in §7/Appendix C. For a paper whose central selling point is exact pathwise accounting, either add residual diagnostics for the bandit/EXP4-IX and continuous-OCO identities or soften the “same in every case” language so that the load-bearing claim remains the finite-exp
minor comments (5)
- Algorithm 1 and §3.1: the crossed cells RET-PRESS and LOC-SQRT are defined but only lightly analyzed; a short remark on when a practitioner would prefer them over the main pairings would help.
- Notation: pt (played), qt (generic one-step), qt,η (temperature-indexed) is introduced late; a one-line glossary near the start of §2 would reduce friction.
- §7 / Appendix C: the paper reports extensive numerical residual checks and baseline comparisons, but the main text does not state whether code is released; a reproducibility note would strengthen the empirical claims.
- Table 1 is a useful reading guide but is long; consider moving part of it to the appendix or tightening the “Conventional manifestation” column.
- Typos and polish: occasional double spaces and long sentences in §1 and §8; a light copy-edit pass would improve readability without changing content.
Circularity Check
No significant circularity: central identities are algebraic bookkeeping from standard Gibbs/one-step balances, not fitted or self-referential predictions.
full rationale
The paper's strongest claims (Theorem 2.10 / Eq. 18 for the prior-retempered update; Proposition 3.13 and Corollary 3.10 for the local pressure-target update) are exact pathwise equalities obtained by composing the one-step mixed-coincidence identity (Corollary 2.3) with the Gibbs variational identity (Lemma 2.6) and terminal potential (Lemma 2.7). Q_t(c) is defined directly as the scaled one-round mixability gap φ_t(η_t)/η_t under the played distribution; the cumulative sum η_t Q_t plus explicit drift D_T and terminal B_T(ρ) is then shown by telescoping algebra to equal composite-loss regret. This is definitional accounting, not a free-parameter fit that is later called a prediction. Schedule constants (C, Γ, a_t) are explicit controller choices whose consequences are derived, not hidden parameters that force the identity. Empirical checks confirm the algebra to machine precision (as expected for identities) and do not fit coefficients to data. There are no load-bearing self-citations, uniqueness theorems imported from the author, or ansatzes smuggled via citation that close the derivation. The work is self-contained against its stated domain (finite one-step log-normalizer). Minor renaming of the cumulative cumulant as 'intrinsic time' is presentational, not circular. Score 0 is therefore appropriate.
Assumptions & free parameters
free parameters (3)
- comparator budget Γ =
problem-dependent; dyadic grid Γ_j = 2^j in Thm. 4.7
- square-root schedule constant C =
1/√2 (default)
- pressure target a_t (or information quota β_t) =
instance-dependent; unit-potential a_t=0 is a special case
assumptions (4)
- standard math Relative entropy and Gibbs variational identities for finite (or continuum) exponential families / softmax posteriors.
- domain assumption Predictable positive learning rates η_t and finite one-step log-normalizers at those rates.
- domain assumption Side information, optimism, and partial feedback enter only through composite or estimated losses fed to the same Bayes update.
- ad hoc to paper Comparator-centered low-noise condition E[(c_t(i)-⟨ρ,c_t⟩)^2] ≤ κ_ρ (μ(i)-⟨ρ,μ⟩) for stochastic luckiness.
invented entities (2)
-
Intrinsic time V_T(c) := Σ_t Q_t(c)
independent evidence
-
2×2 adaptive Bayes design space (RET/LOC × SQRT/PRESS)
Cite this review
Pith. "Pith review of Adaptive Bayes exactly tracks information over intrinsic time." pith.science (2026). https://pith.science/paper/KDUNRKBG
@misc{pith2026260708789,
author = {Pith},
title = {Pith review of: Adaptive Bayes exactly tracks information over intrinsic time},
year = {2026},
howpublished = {\url{https://pith.science/paper/KDUNRKBG}},
note = {Machine review of arXiv:2607.08789}
}
read the original abstract
Bayesian and multiplicative-weights updates reweight experts, models, or actions from sequential feedback. We show that the regret of any such update obeys an exact information-accounting identity. On each round, the learner's excess loss to any chosen comparator is the sum of an immediate payment for the uncertainty exposed by the round and a reduction in the information distance from the learner's current weights to the comparator. The cumulative payment defines a pathwise uncertainty clock, the \emph{intrinsic time} of the realized sequence. Summing one-step balances yields two exact adaptive decompositions of cumulative regret, one for each natural way of composing the update across rounds. Because the decompositions are exact rather than upper bounds, favorable stochastic or low-noise regimes appear as self-bounding properties of the realized intrinsic time, not as slack in worst-case analyses. The same calculus covers Hedge, optimistic and side-information variants, continuous priors, boosting, online convex optimization, contextual bandits, and repeated games: the pathwise account is the same in every case.
Figures
Figures from the paper (16 more)
Reference graph
Works this paper leans on
-
[1]
Amit Agarwal, Elad Hazan, Satyen Kale, and Robert E. Schapire. Algorithms for portfolio management based on the newton method. In International Conference on Machine Learning (ICML), 2006. doi: 10.1145/1143844.1143846
-
[2]
Online learning with feedback graphs: Beyond bandits
Noga Alon, Nicolo Cesa-Bianchi, Ofer Dekel, and T omer Koren. Online learning with feedback graphs: Beyond bandits. Conference on Learning Theory (COLT), 2015
2015
-
[3]
The multiplicative weights update method: a meta-algorithm and appli- cations
Sanjeev Arora, Elad Hazan, and Satyen Kale. The multiplicative weights update method: a meta-algorithm and appli- cations. Theory of Computing, 8:121–164, 2012. doi: 10.4086/toc.2012.v008a006
-
[4]
Uci machine learning repository, 2007
Arthur Asuncion and David Newman. Uci machine learning repository, 2007. URL https://archive.ics.uci. edu/ml
2007
-
[5]
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire. The nonstochastic multiarmed bandit problem. SIAM Journal on Computing, 32(1):48–77, 2002. doi: 10.1137/S0097539701398375. Source-bbl-verified on 2026-05-24; lifted from source paper’s bbl at ingest (bibvac-lifted-from=ACBFS02)
-
[6]
Sharp finite-time iterated-logarithm martingale concentration
Akshay Balsubramani. Sharp finite-time iterated-logarithm martingale concentration. arXiv preprint, 2014
2014
-
[7]
From external to internal regret
Avrim Blum and Yishay Mansour. From external to internal regret. Journal of Machine Learning Research , 8:1307– 1324, 2007. URL https://www.jmlr.org/papers/v8/blum07a.html
2007
-
[8]
Olivier Catoni. PAC-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning , volume 56 of Institute of Mathematical Statistics Lecture Notes–Monograph Series . Institute of Mathematical Statistics, Beachwood, OH, 2007. doi: 10.1214/074921707000000391
Show all 79 references
-
[9]
On prediction of individual sequences
Nicolo Cesa-Bianchi and Gábor Lugosi. On prediction of individual sequences. Annals of Statistics, 27(6):1865–1895,
-
[10]
doi: 10.1214/aos/1017939242
-
[11]
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gábor Lugosi. Prediction, Learning, and Games . Cambridge University Press, 2006. doi: 10.1017/CBO9780511546921
2006 doi
-
[12]
Helmbold, Robert E
Nicolo Cesa-Bianchi, Yoav Freund, David Haussler, David P. Helmbold, Robert E. Schapire, and Manfred K. Warmuth. How to use expert advice. Journal of the ACM, 44(3):427–485, 1997. doi: 10.1145/258128.258179
1997 doi
-
[13]
Improved second-order bounds for prediction with expert advice
Nicolo Cesa-Bianchi, Yishay Mansour, and Gilles Stoltz. Improved second-order bounds for prediction with expert advice. Machine Learning, 66(2–3):321–352, 2007
2007
-
[14]
A parameter-free hedging algorithm
Kamalika Chaudhuri, Yoav Freund, and Daniel Hsu. A parameter-free hedging algorithm. In Advances in Neural Information Processing Systems (NeurIPS), 2009
2009
-
[15]
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff. A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations. Annals of Mathematical Statistics, 23(4):493–507, 1952. doi: 10.1214/aoms/1177729330
1952 doi
-
[16]
Online optimization with gradual variations
Chao-Kai Chiang, Tianbao Yang, Chia-Jung Lee, Mehrdad Mahdavi, Chi-Jen Lu, Rong Jin, and Shenghuo Zhu. Online optimization with gradual variations. InConference on Learning Theory (COLT), 2012. URL https://proceedings. mlr.press/v23/chiang12.html. 104
2012
-
[17]
Thomas M. Cover. Universal portfolios. Mathematical Finance, 1(1):1–29, 1991. doi: 10.1111/j.1467-9965.1991. tb00002.x
1991 doi
-
[18]
Cover and Joy A
Thomas M. Cover and Joy A. Thomas. Elements of Information Theory . Wiley-Interscience, 1991. doi: 10.1002/ 0471200611
1991
-
[19]
Combining online learning guarantees
Ashok Cutkosky. Combining online learning guarantees. In Conference on Learning Theory (COLT), 2019
2019
-
[20]
A. P. Dawid. Statistical theory: The prequential approach. Journal of the Royal Statistical Society. Series A , 147(2): 278–292, 1984. doi: 10.2307/2981683
1984 doi
-
[21]
Grünwald, and Wouter M
Steven de Rooij, Tim van Erven, Peter D. Grünwald, and Wouter M. Koolen. Follow the leader if you can, hedge if you must. Journal of Machine Learning Research, 15:1281–1316, 2014
2014
-
[22]
Forecasting electricity consumption by aggregating specialized experts
Marie Devaine, Pierre Gaillard, Yannig Goude, and Gilles Stoltz. Forecasting electricity consumption by aggregating specialized experts. Machine Learning, 90(2):231–260, 2013. doi: 10.1007/s10994-012-5314-7
2013 doi
-
[23]
Adaptive subgradient methods for online learning and stochastic op- timization
John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic op- timization. Journal of Machine Learning Research , 12:2121–2159, 2011. URL https://jmlr.org/papers/v12/ duchi11a.html
2011
-
[24]
Foster and Rakesh V
Dean P. Foster and Rakesh V. Vohra. Asymptotic calibration. Biometrika, 85(2):379–390, 1998. doi: 10.1093/biomet/ 85.2.379
1998 doi
-
[25]
Foster and Alexander Rakhlin
Dylan J. Foster and Alexander Rakhlin. Beyond UCB: Optimal and efficient contextual bandits with regression ora- cles. In International Conference on Machine Learning (ICML), 2020
2020
-
[26]
Open problem: Second order regret bounds based on scaling time
Yoav Freund. Open problem: Second order regret bounds based on scaling time. In Conference on Learning Theory (COLT), 2016. URL https://proceedings.mlr.press/v49/freund16.html
2016
-
[27]
Schapire
Yoav Freund and Robert E. Schapire. Game theory, on-line prediction and boosting. Conference on Computational Learning Theory (COLT), 1996. doi: 10.1145/238061.238163
1996 doi
-
[28]
Schapire
Yoav Freund and Robert E. Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences , 55(1):119–139, 1997. doi: 10.1006/jcss.1997.1504
1997 doi
-
[29]
Schapire
Yoav Freund and Robert E. Schapire. Adaptive game playing using multiplicative weights. Games and Economic Behavior, 29(1–2):79–103, 1999. doi: 10.1006/game.1999.0738
1999 doi
-
[30]
Schapire, Yoram Singer, and Manfred K
Yoav Freund, Robert E. Schapire, Yoram Singer, and Manfred K. Warmuth. Using and combining predictors that specialize. In ACM Symposium on Theory of Computing (STOC), pages 334–343, 1997. doi: 10.1145/258533.258616
1997 doi
-
[31]
Yoav Freund, Nicholas J. A. Harvey, Victor S. Portella, Yabing Qi, and Yu-Xiang Wang. A second order regret bound for NormalHedge. arXiv preprint arXiv:2602.08151, 2026
2026
-
[32]
A second-order bound with excess losses
Pierre Gaillard, Gilles Stoltz, and Tim van Erven. A second-order bound with excess losses. Conference on Learning Theory (COLT), 2014
2014
-
[33]
The KL-UCB algorithm for bounded stochastic bandits and beyond
Aurélien Garivier and Olivier Cappé. The KL-UCB algorithm for bounded stochastic bandits and beyond. In Pro- ceedings of the 24th Annual Conference on Learning Theory (COLT), pages 359–376, 2011
2011
-
[34]
Combining probability distributions: A critique and an annotated bibliography
Christian Genest and James V Zidek. Combining probability distributions: A critique and an annotated bibliography. Statistical Science, 1(1):114–135, 1986. doi: 10.1214/ss/1177013825
1986 doi
-
[35]
A continuous colonel blotto game
Oliver Gross and Robert Wagner. A continuous colonel blotto game. RAND Research Memorandum RM-408 , 1950. URL https://www.rand.org/pubs/research_memoranda/RM408.html
1950
-
[36]
Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it
Peter Grünwald and Thijs van Ommen. Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it. Bayesian Analysis, 12(4):1069–1103, 2017. doi: 10.1214/17-BA1085. arXiv:1412.3730
-
[37]
Grünwald
Peter D. Grünwald. The Minimum Description Length Principle. MIT Press, 2007. URL https://direct.mit.edu/ books/monograph/3813/The-Minimum-Description-Length-Principle . 105
2007
-
[38]
Grünwald
Peter D. Grünwald. The safe bayesian: Learning the learning rate via the mixability gap. In Algorithmic Learning Theory (ALT 2012), volume 7568 of Lecture Notes in Computer Science , pages 169–183. Springer, 2012. doi: 10.1007/ 978-3-642-34106-9_16
2012
-
[39]
David Haussler, Jyrki Kivinen, and Manfred K. Warmuth. Sequential prediction of individual sequences under general loss functions. In IEEE Transactions on Information Theory , volume 44, pages 1906–1925, 1998. doi: 10.1109/18.705569
1906 doi
-
[40]
Extracting certainty from uncertainty: Regret bounded by variation in costs
Elad Hazan and Satyen Kale. Extracting certainty from uncertainty: Regret bounded by variation in costs. In Con- ference on Learning Theory (COLT), 2010. doi: 10.1007/s10994-010-5175-x
2010 doi
-
[41]
Mark Herbster and Manfred K. Warmuth. Tracking the best expert. Machine Learning, 32(2):151–178, 1998. doi: 10.1023/A:1007424614876
1998 doi
-
[42]
Selecting weighting factors in logarithmic opinion pools
T om Heskes. Selecting weighting factors in logarithmic opinion pools. Advances in Neu- ral Information Processing Systems (NeurIPS) , 1997. URL https://papers.nips.cc/paper/ 1413-selecting-weighting-factors-in-logarithmic-opinion-pools
1997
-
[43]
Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon
Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon. Time-uniform, nonparametric, nonasymp- totic confidence sequences. The Annals of Statistics, 49(2), 2021. doi: 10.1214/20-AOS1991
2021 doi
-
[44]
General linear relations between different types of predictive complexity
Yuri Kalnishkan. General linear relations between different types of predictive complexity. Theoretical Computer Science , 271:181–200, 2002. URL https://pure.royalholloway.ac.uk/en/publications/ general-linear-relations-among-different-types-of-predictive-comp-2/
2002
-
[45]
Efficient learning by im- plicit exploration in bandit problems with side observations
T omáš Kocák, Gergely Neu, Michal Valko, and Rémi Munos. Efficient learning by im- plicit exploration in bandit problems with side observations. In Advances in Neural In- formation Processing Systems (NeurIPS) , 2014. URL http://papers.nips.cc/paper/ 5462-efficient-learning-by...
2014
-
[46]
Koolen and Tim van Erven
Wouter M. Koolen and Tim van Erven. Second-order quantile methods for experts and combinatorial games. In Conference on Learning Theory (COLT), 2015
2015
-
[47]
Koolen, Dmitry Adamskiy, and Manfred K
Wouter M. Koolen, Dmitry Adamskiy, and Manfred K. Warmuth. Putting bayes to sleep. In Ad- vances in Neural Information Processing Systems (NeurIPS) , 2012. URL https://papers.nips.cc/paper/ 4557-putting-bayes-to-sleep
2012
-
[48]
Lattimore and A
T. Lattimore and A. György. Mirror descent and the information ratio. In Conference on Learning Theory, 2021. Source-bbl-verified on 2026-05-24; lifted from source paper’s bbl at ingest (bibvac-lifted- from=LattimoreGyorgy21)
2021
-
[49]
Nick Littlestone and Manfred K. Warmuth. The weighted majority algorithm.Inf. Comput., 108(2):212–261, February
-
[50]
doi: 10.1006/inco.1994.1009
ISSN 0890-5401. doi: 10.1006/inco.1994.1009. URL http://dx.doi.org/10.1006/inco.1994.1009
1994 doi
-
[51]
A short note on a variant of the squint algorithm
Haipeng Luo. A short note on a variant of the squint algorithm. arXiv preprint arXiv:2603.03409, 2026
2026
-
[52]
Schapire
Haipeng Luo and Robert E. Schapire. Achieving all with no parameters: Adanormalhedge. Conference on Learning Theory (COLT), 2015
2015
-
[53]
Marinov and Julian Zimmert
T eodor V. Marinov and Julian Zimmert. The pareto frontier of model selection for general contextual bandits. In Advances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[54]
Universal prediction.IEEE Transactions on Information Theory, 44(6):2124–2147, 1998
Neri Merhav and Meir Feder. Universal prediction.IEEE Transactions on Information Theory, 44(6):2124–2147, 1998. doi: 10.1109/18.720534
1998 doi
-
[55]
G. Neu. Explore no more: Improved high-probability regret bounds for non-stochastic bandits. In Advances in Neural Information Processing Systems, pages 3168–3176, 2015. Source-bbl-verified on 2026-05-24; lifted from source paper’s bbl at ingest (bibvac-lifted-from=Neu15). 106
2015
-
[56]
No-regret learning with unbounded losses: The case of logarithmic pooling
Eric Neyman and Tim Roughgarden. No-regret learning with unbounded losses: The case of logarithmic pooling. In Advances in Neural Information Processing Systems (NeurIPS), 2023
2023
-
[57]
On the chi-square and higher-order chi distances for approximatingf -divergences
Frank Nielsen and Richard Nock. On the chi-square and higher-order chi distances for approximatingf -divergences. IEEE Signal Processing Letters, 21(1):10–13, 2014
2014
-
[58]
Coin betting and parameter-free online learning
Francesco Orabona and Dávid Pál. Coin betting and parameter-free online learning. Advances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[59]
Ortega and Daniel A
Pedro A. Ortega and Daniel A. Braun. Generalized thompson sampling for sequential decision-making and causal inference. Complex Adaptive Systems Modeling, 2:2, 2014
2014
-
[60]
Muriel Felipe Pérez-Ortiz and Wouter M. Koolen. Luckiness in multiscale online learning. In Advances in Neu- ral Information Processing Systems (NeurIPS) , 2022. URL https://proceedings.neurips.cc/paper_files/ paper/2022/hash/a0d2345b43e66fa946155c98899dc03b-Abstract-Conference.html
2022
-
[61]
Manning, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. InAdvances in Neural Information Process- ing Systems (NeurIPS), 2023
2023
-
[62]
On equivalence of martingale tail bounds and deterministic regret in- equalities
Alexander Rakhlin and Karthik Sridharan. On equivalence of martingale tail bounds and deterministic regret in- equalities. In Conference on Learning Theory (COLT), 2017
2017
-
[63]
Universal coding, information, prediction, and estimation
Jorma Rissanen. Universal coding, information, prediction, and estimation. IEEE Transactions on Information Theory, 30(4):629–636, 1984. doi: 10.1109/TIT.1984.1056936
1984 doi
-
[64]
A near-optimal best-of-both-worlds algorithm for online learning with feedback graphs
Chloé Rouyer, Dirk van der Hoeven, Nicolò Cesa-Bianchi, and Yevgeny Seldin. A near-optimal best-of-both-worlds algorithm for online learning with feedback graphs. In Advances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[65]
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy. Learning to optimize via posterior sampling. In Mathematics of Operations Research, volume 39, pages 1221–1243, 2014
2014
-
[66]
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy. Learning to optimize via information-directed sampling. Advances in neural information processing systems, 27, 2014
2014
-
[67]
A tutorial on thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen. A tutorial on thompson sampling. Foundations and Trends in Machine Learning, 11(1):1–96, 2018
2018
-
[68]
Schapire and Yoav Freund
Robert E. Schapire and Yoav Freund. Boosting: Foundations and Algorithms. MIT Press, 2012. doi: 10.7551/mitpress/ 8291.001.0001
2012 doi
-
[69]
Game-Theoretic Foundations for Probability and Finance
Glenn Shafer and Vladimir Vovk. Game-Theoretic Foundations for Probability and Finance. Wiley, 2019. doi: 10.1002/ 9781118548035
2019
-
[70]
Yu. M. Shtarkov. Universal sequential coding of single messages. Problems of Information Transmission, 23(3):3–17,
-
[71]
URL https://www.mathnet.ru/eng/ppi811
-
[72]
On general minimax theorems
Maurice Sion. On general minimax theorems. Pacific Journal of Mathematics, 8(1):171–176, 1958. doi: 10.2140/pjm. 1958.8.171
1958 doi
-
[73]
Adaptivity and optimism: An improved exponentiated gradient algorithm
Jacob Steinhardt and Percy Liang. Adaptivity and optimism: An improved exponentiated gradient algorithm. In International Conference on Machine Learning (ICML) , 2014. URL https://proceedings.mlr.press/v32/ steinhardtb14.html
2014
-
[74]
Syrgkanis, A
V. Syrgkanis, A. Agarwal, H. Luo, and R. E. Schapire. Fast convergence of regularized learning in games. In Advances in Neural Information Processing Systems , pages 2989–2997, 2015. Source-bbl-verified on 2026-05-24; lifted from source paper’s bbl at ingest (bibvac-lifted-fro...
2015
-
[75]
Tim van Erven and Wouter M. Koolen. MetaGrad: multiple learning rates in online learning. In Advances in Neural Information Processing Systems (NeurIPS), pages 3666–3674, 2016. 107
2016
-
[76]
A game of prediction with expert advice
Vladimir Vovk. A game of prediction with expert advice. In Journal of Computer and System Sciences , volume 56, pages 153–173, 1998. doi: 10.1006/jcss.1997.1556
1998 doi
-
[77]
Algorithmic Learning in a Random World
Vladimir Vovk, Alex Gammerman, and Glenn Shafer. Algorithmic Learning in a Random World . Springer, 2005. doi: 10.1007/b106715
2005 doi
-
[78]
Zimmert and Y
J. Zimmert and Y. Seldin. Tsallis-inf: An optimal algorithm for stochastic and adversarial bandits. Journal of Machine Learning Research, 22(28):1–49, 2021. Source-bbl-verified on 2026-05-24; lifted from source paper’s bbl at ingest (bibvac-lifted-from=ZimmertSeldin21)
2021
-
[79]
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich. Online convex programming and generalized infinitesimal gradient ascent. In International Con- ference on Machine Learning (ICML), 2003. doi: 10.5555/3041838.3041955. 108
2003 doi
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.