POPO uses bounded importance sampling on positive rollouts and a siamese policy network to achieve implicit negative gradients and stable optimization, matching or exceeding GRPO on math benchmarks such as 36.67% on AIME 2025.
Biometrics bulletin , volume=
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
background 1polarities
background 1representative citing papers
We derive powerful e-processes for sequential stochastic-dominance testing, with power-one guarantees for first- and higher-order dominance.
Re-evaluating four LLM code-efficiency benchmarks with 30-run statistical testing shows 93.89% of 'performant' implementations are indistinguishable from baselines; a multi-agent test-generation framework reveals hidden significant improvements in ~24% of previously non-significant tasks.
Tabular foundation models, used zero-shot, match or beat tuned gradient boosting on average in credit PD and LGD benchmarks, with a larger edge on small datasets.
Robust-normalized phasic SCR peak rate from 25 Hz wrist GSR separates TSST stress from sitting/standing baselines at balanced accuracies 0.82–0.87 in 31 subjects.
Motion-weighted beat aggregation plus perfusion-index-guided R correction yields wrist SpO2 MAE 2.305% at 25 Hz, comparable to 100 Hz on a private 9-subject dataset.
citing papers explorer
-
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients
POPO uses bounded importance sampling on positive rollouts and a siamese policy network to achieve implicit negative gradients and stable optimization, matching or exceeding GRPO on math benchmarks such as 36.67% on AIME 2025.
-
Betting on Bets: Anytime-Valid Tests for Stochastic Dominance
We derive powerful e-processes for sequential stochastic-dominance testing, with power-one guarantees for first- and higher-order dominance.
-
Rethinking Code Performance Benchmarks for LLMs
Re-evaluating four LLM code-efficiency benchmarks with 30-run statistical testing shows 93.89% of 'performant' implementations are indistinguishable from baselines; a multi-agent test-generation framework reveals hidden significant improvements in ~24% of previously non-significant tasks.
-
Foundation Models for Credit Risk Prediction: A Game Changer?
Tabular foundation models, used zero-shot, match or beat tuned gradient boosting on average in credit PD and LGD benchmarks, with a larger edge on small datasets.
-
Unit-Independent Low-Rate Wrist GSR Processing for Stress Detection Using Phasic nSCR Features
Robust-normalized phasic SCR peak rate from 25 Hz wrist GSR separates TSST stress from sitting/standing baselines at balanced accuracies 0.82–0.87 in 31 subjects.
-
Low-Rate Wrist SpO2 Estimation under Micro-Perturbations Using Motion-Aware Beat Selection and Perfusion-Guided Calibration
Motion-weighted beat aggregation plus perfusion-index-guided R correction yields wrist SpO2 MAE 2.305% at 25 Hz, comparable to 100 Hz on a private 9-subject dataset.