Pith. sign in

REVIEW 14 cited by

The Dark Machines Anomaly Score Challenge: Benchmark Data and Model Independent Event Classification for the Large Hadron Collider

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.14027 v3 pith:HU5XUU4M submitted 2021-05-28 hep-ph hep-exphysics.data-anstat.ML

classification hep-phhep-exphysics.data-anstat.ML
keywords anomalybenchmarkchallengedataphysicsalgorithmsanalysisdark
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We describe the outcome of a data challenge conducted as part of the Dark Machines Initiative and the Les Houches 2019 workshop on Physics at TeV colliders. The challenged aims at detecting signals of new physics at the LHC using unsupervised machine learning algorithms. First, we propose how an anomaly score could be implemented to define model-independent signal regions in LHC searches. We define and describe a large benchmark dataset, consisting of >1 Billion simulated LHC events corresponding to $10~\rm{fb}^{-1}$ of proton-proton collisions at a center-of-mass energy of 13 TeV. We then review a wide range of anomaly detection and density estimation algorithms, developed in the context of the data challenge, and we measure their performance in a set of realistic analysis environments. We draw a number of useful conclusions that will aid the development of unsupervised new physics searches during the third run of the LHC, and provide our benchmark dataset for future studies at https://www.phenoMLdata.org. Code to reproduce the analysis is provided at https://github.com/bostdiek/DarkMachines-UnsupervisedChallenge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking Machine Learning Architectures for ttH Multilepton Signal Sensitivity

    hep-ph 2026-07 conditional novelty 7.0 of 10

    A controlled benchmark of six ML classifiers on a new simulated ttH multilepton dataset finds symmetry-constrained graph models (Particle Transformer, LorentzNet) and azimuthal RoPE encoding outperform tabular baselines.

  2. Enhancing anomaly detection with topology-aware autoencoders

    hep-ph 2025-02 conditional novelty 7.0 of 10

    Autoencoders with latent spaces shaped like S^2, S^2×S^2, or RP^2, matched to the phase-space topology of the background, reduce spurious reconstruction errors and give a small but consistent anomaly-detection gain ov...

  3. Generative Amplification with Surrogate Monte Carlo

    hep-ph 2026-08 conditional novelty 6.0 of 10

    An amplitude surrogate trained on a few thousand exact LHC amplitude points statistically outperforms the training data, with largest amplification in sparsely populated kinematic tails of Z+g and Z+4g production.

  4. Explicit or Implicit? Encoding Physics at the Precision Frontier

    hep-ph 2026-03 conditional novelty 6.0 of 10

    On three precision classification tasks — reweighting-based unfolding, likelihood-ratio estimation, and weakly supervised anomaly detection — a Lorentz-equivariant transformer and a pretrained foundation model perform...

  5. Look everywhere effects in anomaly detection

    hep-ph 2025-12 conditional novelty 6.0 of 10

    Weakly supervised anomaly detectors that train and test on the same data produce badly miscalibrated p-values; independent test sets are calibrated but insensitive, while k-fold cross-validation is a workable middle ground.

  6. SparsePixels: Efficient Convolution for Sparse Data on FPGAs

    cs.AR 2025-12 conditional novelty 6.0 of 10

    A fixed-budget sparse-convolution FPGA framework runs CNNs on <=20 of ~4000 pixels, achieving 0.665 us inference for MicroBooNE with a 73x speedup and ~2% AUC loss.

  7. Graph theory inspired anomaly detection at the LHC

    hep-ph 2025-06 conditional novelty 6.0 of 10

    Sparse globally rigid graph representations of jets, combined with roughly 30 reclustered subjets, improve graph autoencoder anomaly detection on the LHC Olympics benchmark.

  8. Search for new physics in final states with semi-visible jets or anomalous signatures using the ATLAS detector

    hep-ex 2025-05 accept novelty 6.0 of 10

    ATLAS finds no sign of semi-visible jets from Z' decays and excludes Z' masses from 2000 to 3200 GeV for invisible fractions between 0.2 and 0.37.

  9. Automatizing the search for mass resonances using BumpNet

    physics.data-an 2025-01 conditional novelty 6.0 of 10

    One trained convolutional network predicts bump significance across mass histograms of different sizes and backgrounds, approaching the accuracy of the ideal likelihood-ratio test.

  10. Optimal Transport Event Representation for Anomaly Detection

    hep-ph 2025-12 conditional novelty 5.0 of 10

    Adding a few optimal-transport-based features to standard jet observables nearly doubles anomaly-detection significance at 0.5% signal injection on LHC Olympics benchmarks.

  11. Generator Based Inference (GBI)

    hep-ph 2025-05 conditional novelty 5.0 of 10

    Generator Based Inference uses data-derived background generators to turn resonant anomaly detection into parameter estimation, reaching 0.1 sigma signal sensitivity on the LHCO benchmark.

  12. HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency

    hep-ph 2025-12 conditional novelty 4.0 of 10

    HEPTAPOD uses LLM agents to drive FeynRules, MadGraph, Pythia, and analysis tools through schema-validated tool calls and run-card templates, demonstrated on a leptoquark signal scan.

  13. The Living Guide of Machine Learning for Particle Physics

    hep-ph 2026-08 conditional novelty 3.0 of 10

    The HEP-ML Living Review is frozen and replaced by a curated, annotated Living Guide designed for orientation in a mature field.

  14. Machine Learning Power Week 2023: Clustering in Hadronic Calorimeters

    nucl-ex 2025-08 conditional novelty 3.0 of 10

    Seven student teams applied K-means, anti-kt, and graph-based methods to ePIC calorimeter clustering; all beat the benchmark, with K-means variants on spherical coordinates performing best.

Pith tools