Pith. sign in

Paper Citation Record · LEDGER

Multimodal semantic retrieval for product search

As of 11 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2501.07365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07365 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:46:16.272642Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87bd7a2d-ef47-455e-9fc5-68ddf573c684 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Multimodal semantic retrieval for product search Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.194963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.194963Z digest=sha256:a623a9301e52594d54af4996a71615b864caa1eaf13a43a1b75095dc4cd41623

Observation da9b08f8-d38e-4620-96de-4dcbeec87b2a · outbound

This paper cites Extreme multi-label learning for semantic matching in product search.

Multimodal semantic retrieval for product search Extreme multi-label learning for semantic matching in product search

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.511983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.199473Z digest=sha256:290e5bdb3cd2b3c0fd3dcd8028e750ace4097e772922c2cba19ddf317f9dc649

Observation c1b8c70d-fc11-4223-83ac-440b211ee79e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Multimodal semantic retrieval for product search A simple framework for contrastive learning of visual representations

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.499154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.203389Z digest=sha256:cee95a7e2186cfb92a6017c13d1ef31aae1e7e3d72b6006e3610fffc6b4aa066

Observation 13307cb4-405f-4ee6-8cba-cb22cba548da · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Multimodal semantic retrieval for product search An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.207475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.207475Z digest=sha256:71a671c48ed832a01e42446c26b64db6b658dd53c247e0583c23e8536f3003d1

Observation 25d3bf16-cb26-4bc2-824e-d762ec0dab1e · outbound

This paper cites Data Filtering Networks.

Multimodal semantic retrieval for product search Data Filtering Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.211886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.211886Z digest=sha256:979b4ec9bf355c7bcaadf2025d56e5ab3bcbb58672e60bb44598fd5d7b0d5d63

Observation 1aae9dbf-a023-4c7a-b1e0-c3db88923259 · outbound

This paper cites Learning deep structured semantic models for web search using clickthrough data.

Multimodal semantic retrieval for product search Learning deep structured semantic models for web search using clickthrough data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.486530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.216147Z digest=sha256:a82bf2370f6c49ce8e1145c8702493523db73da38ad6095a52a879190c31bcc0

Observation a6a95a11-f280-42de-8ed1-84701725b540 · outbound

This paper cites Bert: Pre- training of deep bidirectional transformers for language understanding.

Multimodal semantic retrieval for product search Bert: Pre- training of deep bidirectional transformers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.473588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.220141Z digest=sha256:243e4c21e872d3a29498f9b3f0e52c9750c9f652c48837ee54726d77a9b8e77d

Observation 083b94a2-312b-4c4d-808e-d54034a2d0df · outbound

This paper cites Relevance-guided supervi- sion for openqa with colbert.

Multimodal semantic retrieval for product search Relevance-guided supervi- sion for openqa with colbert

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.461330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.223612Z digest=sha256:222e2cdd11182e0cf8e496e9574aee420fafdb15a470405b6eba57e27002c6e6

Observation f25b13c2-3197-4bb3-819d-70570b8545df · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

Multimodal semantic retrieval for product search Vilt: Vision-and-language transformer without convolution or region supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.227144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.227144Z digest=sha256:1466560532b665265d927600a1b7c8795e5dd142f0b1f1da0af61c96418f51f9

Observation ebe36528-50ff-495e-94c7-8dbab2fbd069 · outbound

This paper cites Embracing Structure in Data for Billion-Scale Semantic Product Search.

Multimodal semantic retrieval for product search Embracing Structure in Data for Billion-Scale Semantic Product Search

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.230595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.230595Z digest=sha256:9946af8ef266f091ce34f19bcafc147a5cc0ed8714250702308898f10332a1b8

Observation 86850806-4794-49fc-9955-0b4b98ea031d · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and gen- eration.

Multimodal semantic retrieval for product search BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and gen- eration

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.441106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.234744Z digest=sha256:0193f0a275686bd6bb59db2a1476ffcc1632256643aeaa1654b5dda19026caad

Observation 8001befc-39be-49df-b6c9-b9a088a1e7af · outbound

This paper cites Embedding-based product retrieval in taobao search.

Multimodal semantic retrieval for product search Embedding-based product retrieval in taobao search

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.430600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.238434Z digest=sha256:f52c455bb84304418e36538fc674e537613299e10315b1503ddce3cddf6912a9

Observation a2f4c79a-eda3-4a11-8c59-2521b90486fd · outbound

This paper cites Deep self-adaptive hashing for image retrieval.

Multimodal semantic retrieval for product search Deep self-adaptive hashing for image retrieval

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.419771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.242178Z digest=sha256:32122afcff4e3d8f0ee15e143e076932d6f323b556203a9747585973b6e22bb0

Observation 02e29bcf-64c3-4a95-ab97-dd2103fbad97 · outbound

This paper cites One picture is worth a thousand words? the pricing power of images in e-commerce.

Multimodal semantic retrieval for product search One picture is worth a thousand words? the pricing power of images in e-commerce

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.407681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.245939Z digest=sha256:6c07d2d11ff51d290fbb53f363d77b2142ed19645986449e68e29d9b7f81daa3

Observation 79336b39-1048-4115-be30-0bbf3813f663 · outbound

This paper cites Dc-bert: Decoupling question and document for efficient contextual encoding.

Multimodal semantic retrieval for product search Dc-bert: Decoupling question and document for efficient contextual encoding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.395179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.249501Z digest=sha256:d776d138c8d5a033e7d5dafdd25285e1e397caeb7804ff2ab43a11e7a6649798

Observation 0bc18342-8fe9-4e35-abde-c93e351e474a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal semantic retrieval for product search Learning transferable visual models from natural language supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.253172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.253172Z digest=sha256:e2bba63235909f6a6b496db5683bc508876ca88485478857f31de95d288a6b8c

Observation 17e39bd2-40ae-4e50-96fc-609d8e7e6dc5 · outbound

This paper cites Introduction to information retrieval, volume 39.

Multimodal semantic retrieval for product search Introduction to information retrieval, volume 39

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.256947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.256947Z digest=sha256:e40815fe6cecec9e83eb4c25fdb86d488a59d3f5812ba00037b5dd3b4321aded

Observation 0047cd88-5228-4bd7-83ab-c9c093c0b4d8 · outbound

This paper cites RepBERT: Contextualized Text Embeddings for First-Stage Retrieval.

Multimodal semantic retrieval for product search RepBERT: Contextualized Text Embeddings for First-Stage Retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.260174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.260174Z digest=sha256:bce7ebf93208f89964e95bfbf2d5a8942cf6905cb535f51ceaf3cd544ceb2b03

Observation fcce507d-7469-486a-99cc-532d1796031f · outbound

This paper cites Binary neural network hashing for image retrieval.

Multimodal semantic retrieval for product search Binary neural network hashing for image retrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.368497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.263965Z digest=sha256:09c3246076ea0f4f74bb0cfd0be484ebba415f9f25586ea82838099f0a73a910

Observation 3ae9d517-66e5-4be6-93ed-e85810624113 · outbound

This paper cites Bringing multimodality to amazon visual search system.

Multimodal semantic retrieval for product search Bringing multimodality to amazon visual search system

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.357561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.268023Z digest=sha256:a652b49425adfbee28b296a3f0e38a54c002f961328294e9ad421301bf181d92

Observation cb9e5e22-5576-40eb-9b4d-dc6446244ec3 · outbound

This paper cites Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks.

Multimodal semantic retrieval for product search Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.346018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:46:16.272642Z digest=sha256:12d115f8551b8c5d38259da056db4cd4bead35762de8e9b06544f71ac5ba0fd5

Pith citing papers

No inbound Pith citation observations are available.