REVIEW 2 cited by
Asymptotically Optimal Knockoff Statistics via the Masked Likelihood Ratio
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Asymptotically Optimal Knockoff Statistics via the Masked Likelihood Ratio
read the original abstract
In feature selection problems, knockoffs are synthetic controls for the original features. Employing knockoffs allows analysts to use nearly any variable importance measure or "feature statistic" to select features while rigorously controlling false positives. However, it is not clear which statistic maximizes power. In this paper, we argue that state-of-the-art lasso-based feature statistics often prioritize features that are unlikely to be discovered, leading to low power in real applications. Instead, we introduce masked likelihood ratio (MLR) statistics, which prioritize features according to one's ability to distinguish each feature from its knockoff. Although no single feature statistic is uniformly most powerful in all situations, we show that MLR statistics asymptotically maximize the number of discoveries under a user-specified Bayesian model of the data. (Like all feature statistics, MLR statistics always provide frequentist error control.) This result places no restrictions on the problem dimensions and makes no parametric assumptions; instead, we require a "local dependence" condition that depends only on known quantities. In simulations and three real applications, MLR statistics outperform state-of-the-art feature statistics, including in settings where the Bayesian model is misspecified. We implement MLR statistics in the python package knockpy; our implementation is often faster than computing a cross-validated lasso.
Forward citations
Cited by 2 Pith papers
-
PRADAS: PRior-Assisted DAta Splitting for False Discovery Rate Control
PRADAS derives a Bayes-optimal mirror statistic for any splitting scheme, establishes asymptotic FDR control under weak dependence, and optimizes the split ratio as a stopping time to improve power over standard equal...
-
Estimating the local false discovery rate under an unknown symmetric null
Proposes estimating lfdr via the surrogate density ratio f(-w)/f(w) using logistic regression with natural cubic splines, with asymptotic control if the estimator is consistent.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.