Pith. sign in

REVIEW 2 cited by

A New Unbiased and Efficient Class of LSH-Based Samplers and Estimators for Partition Function Computation in Log-Linear Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1703.05160 v1 pith:OJXTZ7KA submitted 2017-03-15 stat.ML cs.DBcs.DScs.LG

classification stat.MLcs.DBcs.DScs.LG
keywords modelssamplingefficientfunctionpartitionaccuratelyclassestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Log-linear models are arguably the most successful class of graphical models for large-scale applications because of their simplicity and tractability. Learning and inference with these models require calculating the partition function, which is a major bottleneck and intractable for large state spaces. Importance Sampling (IS) and MCMC-based approaches are lucrative. However, the condition of having a "good" proposal distribution is often not satisfied in practice. In this paper, we add a new dimension to efficient estimation via sampling. We propose a new sampling scheme and an unbiased estimator that estimates the partition function accurately in sub-linear time. Our samples are generated in near-constant time using locality sensitive hashing (LSH), and so are correlated and unnormalized. We demonstrate the effectiveness of our proposed approach by comparing the accuracy and speed of estimating the partition function against other state-of-the-art estimation techniques including IS and the efficient variant of Gumbel-Max sampling. With our efficient sampling scheme, we accurately train real-world language models using only 1-2% of computations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval

    cs.IR 2026-07 conditional novelty 6.0 of 10

    OneShot trains hierarchical item codebooks jointly with the ranking loss, enabling nonlinear neural scoring in billion-scale retrieval and reporting +20% recall, 10x fewer dense-ranked items, and live Instagram gains.

  2. Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever

    cs.IR 2025-01 conditional novelty 6.0 of 10

    A streaming vector-quantization index that updates item-cluster assignments in real time outperforms and replaces HNSW, Deep Retrieval, and NANN retrievers in Douyin and Douyin Lite production.

Pith tools