Pith. sign in

REVIEW 6 cited by

Energy-based Automated Model Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12689 v3 pith:Q5BTZPWD submitted 2024-01-23 cs.LG cs.AIcs.CLcs.CV

classification cs.LGcs.AIcs.CLcs.CV
keywords autoevalenergyevaluationlearningautomatedenergy-basedlabelsmeta-distribution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The conventional evaluation protocols on machine learning models rely heavily on a labeled, i.i.d-assumed testing dataset, which is not often present in real world applications. The Automated Model Evaluation (AutoEval) shows an alternative to this traditional workflow, by forming a proximal prediction pipeline of the testing performance without the presence of ground-truth labels. Despite its recent successes, the AutoEval frameworks still suffer from an overconfidence issue, substantial storage and computational cost. In that regard, we propose a novel measure -- Meta-Distribution Energy (MDE) -- that allows the AutoEval framework to be both more efficient and effective. The core of the MDE is to establish a meta-distribution statistic, on the information (energy) associated with individual samples, then offer a smoother representation enabled by energy-based learning. We further provide our theoretical insights by connecting the MDE with the classification loss. We provide extensive experiments across modalities, datasets and different architectural backbones to validate MDE's validity, together with its superiority compared with prior approaches. We also prove MDE's versatility by showing its seamless integration with large-scale models, and easy adaption to learning scenarios with noisy- or imbalanced- labels. Code and data are available: https://github.com/pengr/Energy_AutoEval

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Co-Adaptive Multi-Task LoRA: Transfer-Aware, Label-Free Control of Domain Participation

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A forward-only controller sets multi-domain LoRA participation from label-free competence and cross-domain affinity, improving average accuracy while using half the data.

  2. Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Replacing fixed top-k routing in MoE-LoRA with router-confidence-based nucleus admission plus an expert-disagreement extension improves accuracy and OOD detection at matched average compute.

  3. Online Data Selection Is Implicit Alignment

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Online SFT data selection acts as an implicit preference model, shifting refusal rates, verbosity, and sycophancy in directions predictable from the selected data's attribute mixture.

  4. Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A statistical non-inferiority test on estimated per-sample correctness probabilities flags when a classifier's accuracy on unlabeled user data drops by more than a chosen margin relative to its test set.

  5. Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A training-dynamics abstention method matches deep ensembles at a fraction of the training cost, and a five-term error budget explains why selective classifiers still fall short of the oracle.

  6. Towards Unsupervised Model Selection for Domain Adaptive Object Detection

    cs.CV 2024-12 conditional novelty 5.0 of 10

    DAS combines a flatness index score and a prototype distance ratio to select near-optimal DAOD checkpoints without target labels.

Pith tools