Pith. sign in

REVIEW 1 cited by

Embracing Structure in Data for Billion-Scale Semantic Product Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.06125 v1 pith:E64IWHQV submitted 2021-10-12 cs.IR cs.LG

classification cs.IRcs.LG
keywords searchdyadicembeddinggiveninferenceproducttrainingbillion-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present principled approaches to train and deploy dyadic neural embedding models at the billion scale, focusing our investigation on the application of semantic product search. When training a dyadic model, one seeks to embed two different types of entities (e.g., queries and documents or users and movies) in a common vector space such that pairs with high relevance are positioned nearby. During inference, given an embedding of one type (e.g., a query or a user), one seeks to retrieve the entities of the other type (e.g., documents or movies, respectively) that are highly relevant. In this work, we show that exploiting the natural structure of real-world datasets helps address both challenges efficiently. Specifically, we model dyadic data as a bipartite graph with edges between pairs with positive associations. We then propose to partition this network into semantically coherent clusters and thus reduce our search space by focusing on a small subset of these partitions for a given input. During training, this technique enables us to efficiently mine hard negative examples while, at inference, we can quickly find the nearest neighbors for a given embedding. We provide offline experimental results that demonstrate the efficacy of our techniques for both training and inference on a billion-scale Amazon.com product search dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Multi-field Representations for Two-Stage E-commerce Retrieval

    cs.IR 2025-01 conditional novelty 5.0 of 10

    CHARM encodes product fields into a hierarchy of embeddings with block-triangular attention and combines coarse shortlisting with fine field-level reranking, showing modest gains on three e-commerce retrieval benchmarks.

Pith tools