Pith. sign in

REVIEW 1 cited by

Feature Selection with Distance Correlation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.00046 v1 pith:NUASMSM6 submitted 2022-11-30 hep-ph cs.LGhep-exphysics.data-an

classification hep-phcs.LGhep-exphysics.data-an
keywords featurefeaturesselectioncorrelationdistancemanymethodtasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Choosing which properties of the data to use as input to multivariate decision algorithms -- a.k.a. feature selection -- is an important step in solving any problem with machine learning. While there is a clear trend towards training sophisticated deep networks on large numbers of relatively unprocessed inputs (so-called automated feature engineering), for many tasks in physics, sets of theoretically well-motivated and well-understood features already exist. Working with such features can bring many benefits, including greater interpretability, reduced training and run time, and enhanced stability and robustness. We develop a new feature selection method based on Distance Correlation (DisCo), and demonstrate its effectiveness on the tasks of boosted top- and $W$-tagging. Using our method to select features from a set of over 7,000 energy flow polynomials, we show that we can match the performance of much deeper architectures, by using only ten features and two orders-of-magnitude fewer model parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Step Toward Interpretability: Smearing the Likelihood

    hep-ph 2025-01 conditional novelty 6.0 of 10

    Smearing the likelihood over an energy metric reveals the physical scales used by a jet classifier, and the needed smearing radius follows a power-law scaling with dataset size.

Pith tools