REVIEW 6 cited by
A Survey on Evaluation of Out-of-Distribution Generalization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Machine learning models, while progressively advanced, rely heavily on the IID assumption, which is often unfulfilled in practice due to inevitable distribution shifts. This renders them susceptible and untrustworthy for deployment in risk-sensitive applications. Such a significant problem has consequently spawned various branches of works dedicated to developing algorithms capable of Out-of-Distribution (OOD) generalization. Despite these efforts, much less attention has been paid to the evaluation of OOD generalization, which is also a complex and fundamental problem. Its goal is not only to assess whether a model's OOD generalization capability is strong or not, but also to evaluate where a model generalizes well or poorly. This entails characterizing the types of distribution shifts that a model can effectively address, and identifying the safe and risky input regions given a model. This paper serves as the first effort to conduct a comprehensive review of OOD evaluation. We categorize existing research into three paradigms: OOD performance testing, OOD performance prediction, and OOD intrinsic property characterization, according to the availability of test data. Additionally, we briefly discuss OOD evaluation in the context of pretrained models. In closing, we propose several promising directions for future research in OOD evaluation.
Forward citations
Cited by 6 Pith papers
-
Quantization Meets OOD: Generalizable Quantization-aware Training from a Flatness Perspective
Quantization-aware training degrades out-of-distribution accuracy, and a flatness-aware method with gradient-disorder freezing, FQAT, partially recovers it.
-
Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms
Pretraining on mixed ECG cohorts improves in-distribution accuracy but hurts out-of-distribution transfer, and sampling single-cohort batches during pretraining mitigates this.
-
Incremental Uncertainty-aware Performance Monitoring with Active Labeling Intervention
IUPM combines incremental optimal transport matching with uncertainty quantification to estimate deployed model accuracy under gradual shifts and to trigger active labeling only when needed.
-
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks
VEGA ranks pre-trained vision-language models for unlabeled datasets by measuring how well a model's own pseudo-labeled visual clusters align with class-name text features.
-
GAQAT: gradient-adaptive quantization-aware training for domain generalization
A gradient-disorder trigger that selectively freezes task gradients of quantizer scale factors improves quantized domain-generalization accuracy, including near-lossless 4-bit results on DomainNet.
-
Boosting Test Performance with Importance Sampling--a Subpopulation Perspective
The paper derives a closed-form importance weight for subpopulation shift, but the main formula contradicts the text and seems to mis-weight minority samples, and the claimed SOTA worst-group performance does not hold...
Discussion (0). Continue with ORCID to comment.