Pith. sign in

REVIEW 5 cited by

Batch Active Learning Using Determinantal Point Processes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.07975 v1 pith:QTRMC44K submitted 2019-06-19 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningactivebatchdatasamplescomputationalmethodspoint
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data collection and labeling is one of the main challenges in employing machine learning algorithms in a variety of real-world applications with limited data. While active learning methods attempt to tackle this issue by labeling only the data samples that give high information, they generally suffer from large computational costs and are impractical in settings where data can be collected in parallel. Batch active learning methods attempt to overcome this computational burden by querying batches of samples at a time. To avoid redundancy between samples, previous works rely on some ad hoc combination of sample quality and diversity. In this paper, we present a new principled batch active learning method using Determinantal Point Processes, a repulsive point process that enables generating diverse batches of samples. We develop tractable algorithms to approximate the mode of a DPP distribution, and provide theoretical guarantees on the degree of approximation. We further demonstrate that an iterative greedy method for DPP maximization, which has lower computational costs but worse theoretical guarantees, still gives competitive results for batch active learning. Our experiments show the value of our methods on several datasets against state-of-the-art baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SeDi-Instruct: Enhancing Alignment of Language Models through Self-Directed Instruction Generation

    cs.CL 2025-02 reject novelty 6.0 of 10

    SeDi-Instruct generates instruction data by relaxing duplicate filtering, sampling cluster-balanced batches, and replacing low-scoring seeds with instructions from high-gradient-norm batches.

  2. Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning

    eess.AS 2026-07 conditional novelty 5.0 of 10

    BADGE-Greedy-DPP selects active-learning batches by greedily maximizing the volume of gradient embeddings, improving rare-call-type discovery in long-tailed frame-level bioacoustic classification.

  3. Variational Proximal Policy Optimization

    stat.ML 2026-06 unverdicted novelty 5.0 of 10

    VP2O maps PPO to SVGD in a MoE architecture using functional kernels and expert orthogonalization, claiming +179 ELO on Codeforces and 32% token reduction on AIME for a 33B/4B model.

  4. Active learning for photonic crystals

    physics.optics 2026-01 unverdicted novelty 5.0 of 10

    Analytic LL-BNN active learning achieves up to 2.7x reduction in training data for band gap prediction in 2D two-tone photonic crystals while maintaining accuracy.

  5. Determinantal point process sampling for bioacoustic active learning

    cs.SD 2026-07 conditional novelty 4.0 of 10

    CARE-DPP combines class-balanced uncertainty, annealed embedding novelty, and DPP-based batch diversification for bioacoustic active learning, achieving 0.50 mean AULC versus 0.46 for CoreSet.

Pith tools