REVIEW 18 cited by
Long-tail learning via logit adjustment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Long-tail learning via logit adjustment
read the original abstract
Real-world classification problems typically exhibit an imbalanced or long-tailed label distribution, wherein many labels are associated with only a few samples. This poses a challenge for generalisation on such labels, and also makes na\"ive learning biased towards dominant labels. In this paper, we present two simple modifications of standard softmax cross-entropy training to cope with these challenges. Our techniques revisit the classic idea of logit adjustment based on the label frequencies, either applied post-hoc to a trained model, or enforced in the loss during training. Such adjustment encourages a large relative margin between logits of rare versus dominant labels. These techniques unify and generalise several recent proposals in the literature, while possessing firmer statistical grounding and empirical performance.
Forward citations
Cited by 18 Pith papers
-
From Perturbation Correction to Geometry-Aware Sampling: Sharpness-Guided Equilibrium Sampling for Balanced Flat Minima in Long-Tailed Learning
Sharpness-Guided Equilibrium Sampling reweights long-tailed training batches using cumulative class counts and SAM perturbation-loss gaps, improving tail accuracy by up to 10.8 points.
-
Loss Landscape Topology Reveals Why Simple Baselines are Competitive at 3D Point Cloud Segmentation Under Class Imbalance
Uniform cross-entropy is competitive with 11 imbalance-aware losses in point-based 3D segmentation; specialized methods give small, architecture-dependent gains.
-
Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability
A dual-query scene graph generation method unifies detector-based and query-based reasoning in a single decoder, achieving state-of-the-art results on Visual Genome, Open Images v6, and GQA-200.
-
Class-frequency Guided Noise Schedule for Diffusion Models
Proposes CFRG noise schedule for diffusion models that assigns larger noises to low-frequency classes to improve generation on imbalanced datasets.
-
Modular Diffusion Models for Structured Visual Recognition
Modular Diffusion Models decompose diffusion into task-specific modules to model distributions over structured visual outputs for detection, segmentation, and scene graph generation.
-
Rethinking Loss Reweighting for Imbalance Learning as an Inverse Problem: A Neural Collapse Point of View
Loss reweighting is cast as an inverse problem that dynamically infers class weights to equalize per-class average losses under the Neural Collapse simplex ETF target.
-
Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior Adjustment
DistPFN is a test-time posterior adjustment that rescales TabPFN class probabilities to reduce overfitting to the training class distribution under label shift.
-
Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior Adjustment
DistPFN is a test-time posterior adjustment technique that mitigates label shift in TabPFN by downweighting the training prior and emphasizing the model's predicted posterior, with a temperature-scaled variant, evalua...
-
SPECTRA: Spectral Domain-Aware Graph Generation for Imbalanced Molecular Property Regression
SPECTRA improves molecular property regression on underrepresented targets via spectral graph generation with rarity-aware budgeting and Laplacian interpolation, paired with edge-aware Chebyshev GNNs, yielding competi...
-
LoFT: Parameter-Efficient Fine-Tuning for Long-tailed Semi-Supervised Learning in Open-World Scenarios
LoFT uses parameter-efficient fine-tuning of foundation models for long-tailed semi-supervised learning, supported by proofs that this reduces hypothesis complexity to minimize balanced posterior error and compresses ...
-
Mutually Exclusive Multiclass Lesion Segmentation in Neuroimaging: Binary-Guided Weak Supervision with Inter-Class Orthogonality
Binary-guided mutual exclusivity with inter-class orthogonality yields accurate multiclass weakly supervised neuroimaging lesion segmentation from image-level labels alone.
-
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
A multi-modal extension of multi-expert architectures uses confidence-guided fusion from modality-specific networks to handle long-tailed class imbalance across heterogeneous inputs.
-
STaR-DRO: Stateful Tsallis Reweighting for Group-Robust Structured Prediction
STaR-DRO applies momentum-smoothed Tsallis reweighting to focus learning on hard groups in structured prediction, yielding F1 gains on clinical label extraction.
-
SciLT: Long-tailed Image Classification under Scientific Image Domains
On scientific long-tailed image tasks, foundation-model fine-tuning gains are limited; SciLT fuses penultimate and final ViT features under dual supervision to improve balanced accuracy.
-
Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
TaxoNet uses a dual-margin objective to reshape decision boundaries in long-tailed fine-grained plant taxonomy, improving rare-class geometry under open-world conditions.
-
Automatic Dataset Construction (ADC): Sample Collection, Data Curation, and Beyond
The ADC method automates the creation of large image classification datasets using LLMs and search engines, achieving 79% human agreement and reducing label noise on a 1 million image clothing dataset, while also rele...
-
Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention
MSLA is a new attention mechanism that models multi-scale and cross-layer interactions to achieve more accurate OBI recognition than prior attention methods.
-
Dynamic Distillation and Gradient Consistency for Robust Long-Tailed Incremental Learning
Gradient consistency regularization and entropy-driven dynamic distillation improve accuracy by up to 5% in long-tailed incremental learning, with strong gains in majority-to-minority task ordering.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.