REVIEW 1 cited by
Boosted top tagging and its interpretation using Shapley values
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Top tagging has emerged as a fast-evolving subject due to the top quark's significant role in probing physics beyond the standard model. For the reconstruction of top jets, machine learning models have shown a substantial improvement in the classification performance compared to the previous methods. In this work, we build top taggers using $N$-Subjettiness ratios and several Energy Correlation observables as input features to train the eXtreme Gradient BOOSTed decision tree (XGBOOST). The study finds that tighter parton-level matching lead to more accurate tagging. However, in real experimental data, where the parton level data are unknown, this matching cannot be done. We train the XGBOOST models without performing this matching and show that this difference impacts the taggers' effectiveness. Additionally, we test the tagger under different simulation conditions, including changes in center-of-mass energy, parton distribution functions (PDFs), and pileup effects, demonstrating its robustness with performance deviations of less than 1%. Furthermore, we use the SHapley Additive exPlanation (SHAP) framework to calculate the importance of the features of the trained models. It helps us to estimate how much each feature of the data contributed to the model's prediction and what regions are of more importance for each input variable. Finally, we combine all the tagger variables to form a hybrid tagger and interpret the results using the Shapley values.
Forward citations
Cited by 1 Pith paper
-
A Step Toward Interpretability: Smearing the Likelihood
Smearing the likelihood over an energy metric reveals the physical scales used by a jet classifier, and the needed smearing radius follows a power-law scaling with dataset size.
Discussion (0). Continue with ORCID to comment.