Pith. sign in

REVIEW 2 cited by

Variable Selection for Kernel Two-Sample Tests

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.07415 v4 pith:WMV5NFZE submitted 2023-02-15 stat.ML stat.AP

classification stat.MLstat.AP
keywords kernelselectionvariablevariablesderiveformulationsframeworkkernels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider the variable selection problem for two-sample tests, aiming to select the most informative variables to determine whether two collections of samples follow the same distribution. To address this, we propose a novel framework based on the kernel maximum mean discrepancy (MMD). Our approach seeks a subset of variables with a pre-specified size that maximizes the variance-regularized kernel MMD statistic. We focus on three commonly used types of kernels: linear, quadratic, and Gaussian. From a computational perspective, we derive mixed-integer programming formulations and propose exact and approximation algorithms with performance guarantees to solve these formulations. From a statistical viewpoint, we derive the rate of testing power of our framework under appropriate conditions. These results show that the sample size requirements for the three kernels depend crucially on the number of selected variables, rather than the data dimension. Experimental results on synthetic and real datasets demonstrate the superior performance of our method, compared to other variable selection frameworks, particularly in high-dimensional settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Post Hoc Inference for Component Attribution in Multivariate Change-Point Detection

    stat.ME 2026-07 conditional novelty 6.0 of 10

    After a multivariate change-point is detected, a grid-based or sample-splitting two-sample test determines whether a pre-specified block of coordinates changed, with Type I error bounded by α0+α1.

  2. Word Sense Detection Leveraging Maximum Mean Discrepancy

    cs.CL 2025-06 conditional novelty 5.0 of 10

    An MMD-based variable selection method ranks words whose embedding dimensions shifted most between time periods, with qualitative evidence from Japanese news and American English.

Pith tools