Pith. sign in

REVIEW 2 cited by

Shifts 2.0: Extending The Dataset of Real Distributional Shifts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.15407 v2 pith:INGKD4TJ submitted 2022-06-30 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords tasksshiftsdatadatasetbenchmarksdatasetsdistributionalallow
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Distributional shift, or the mismatch between training and deployment data, is a significant obstacle to the usage of machine learning in high-stakes industrial applications, such as autonomous driving and medicine. This creates a need to be able to assess how robustly ML models generalize as well as the quality of their uncertainty estimates. Standard ML baseline datasets do not allow these properties to be assessed, as the training, validation and test data are often identically distributed. Recently, a range of dedicated benchmarks have appeared, featuring both distributionally matched and shifted data. Among these benchmarks, the Shifts dataset stands out in terms of the diversity of tasks as well as the data modalities it features. While most of the benchmarks are heavily dominated by 2D image classification tasks, Shifts contains tabular weather forecasting, machine translation, and vehicle motion prediction tasks. This enables the robustness properties of models to be assessed on a diverse set of industrial-scale tasks and either universal or directly applicable task-specific conclusions to be reached. In this paper, we extend the Shifts Dataset with two datasets sourced from industrial, high-risk applications of high societal importance. Specifically, we consider the tasks of segmentation of white matter Multiple Sclerosis lesions in 3D magnetic resonance brain images and the estimation of power consumption in marine cargo vessels. Both tasks feature ubiquitous distributional shifts and a strict safety requirement due to the high cost of errors. These new datasets will allow researchers to further explore robust generalization and uncertainty estimation in new situations. In this work, we provide a description of the dataset and baseline results for both tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ConfLUNet: Multiple sclerosis lesion instance segmentation in presence of confluent lesions

    eess.IV 2025-05 conditional novelty 6.0 of 10

    ConfLUNet, an end-to-end instance segmentation network for multiple sclerosis lesions, improves separation and counting of confluent lesion units compared with connected components and automatic splitting.

  2. F3-Net: Foundation Model for Full Abnormality Segmentation of Medical Images with Flexible Input Modality Requirement

    cs.CV 2025-07 reject novelty 3.0 of 10

    F3-Net combines multi-encoder nnU-Net with zero-filled missing modalities to segment glioma, metastasis, stroke, and white matter lesions, but the missing-modality claim is untested and comparisons are incomplete.

Pith tools