Pith. sign in

REVIEW 2 cited by

Rethinking InfoNCE: How Many Negative Samples Do You Need?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.13003 v1 pith:AK4AI6LH submitted 2021-05-27 cs.LG cs.IR

classification cs.LGcs.IR
keywords negativesamplingtrainingmodelratiosamplesinfoncemany
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

InfoNCE loss is a widely used loss function for contrastive model training. It aims to estimate the mutual information between a pair of variables by discriminating between each positive pair and its associated $K$ negative pairs. It is proved that when the sample labels are clean, the lower bound of mutual information estimation is tighter when more negative samples are incorporated, which usually yields better model performance. However, in many real-world tasks the labels often contain noise, and incorporating too many noisy negative samples for model training may be suboptimal. In this paper, we study how many negative samples are optimal for InfoNCE in different scenarios via a semi-quantitative theoretical framework. More specifically, we first propose a probabilistic model to analyze the influence of the negative sampling ratio $K$ on training sample informativeness. Then, we design a training effectiveness function to measure the overall influence of training samples on model learning based on their informativeness. We estimate the optimal negative sampling ratio using the $K$ value that maximizes the training effectiveness function. Based on our framework, we further propose an adaptive negative sampling method that can dynamically adjust the negative sampling ratio to improve InfoNCE based model training. Extensive experiments on different real-world datasets show our framework can accurately predict the optimal negative sampling ratio in different tasks, and our proposed adaptive negative sampling method can achieve better performance than the commonly used fixed negative sampling ratio strategy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection

    cs.CR 2025-08 conditional novelty 5.0 of 10

    FALCON automates the generation of Snort and YARA intrusion detection rules from cyber threat intelligence using an LLM agent pipeline with a contrastively trained CTI-rule semantic scorer as a ground-truth-free validator.

  2. TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    TRUST adapts a vision model to an unlabeled target domain by generating pseudo-labels from captions, weighting them by caption-based uncertainty, and aligning image and text features with a soft contrastive loss, repo...

Pith tools