Pith. sign in

Paper Citation Record · LEDGER

Multimodal semantic retrieval for product search

As of 22 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2501.07365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07365 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:46:16.272642Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 87bd7a2d-ef47-455e-9fc5-68ddf573c684 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Multimodal semantic retrieval for product search Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.194963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.194963Z digest=sha256:5f28bb34697aaa4f38f28fa69df0a0866e7575be2a66c548d25fce13054a8ba0

Observation da9b08f8-d38e-4620-96de-4dcbeec87b2a · outbound

This paper cites Extreme multi-label learning for semantic matching in product search.

Multimodal semantic retrieval for product search Extreme multi-label learning for semantic matching in product search

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.511983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.199473Z digest=sha256:7b89e47a3c9804f7c1242cded6c4a3065145fa06244d4873168e1a6066368671

Observation c1b8c70d-fc11-4223-83ac-440b211ee79e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Multimodal semantic retrieval for product search A simple framework for contrastive learning of visual representations

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.499154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.203389Z digest=sha256:d9f383c588792f30153b84286cd9b0c095861b01cb6c51b8f5b9371a48691361

Observation 13307cb4-405f-4ee6-8cba-cb22cba548da · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Multimodal semantic retrieval for product search An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.207475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.207475Z digest=sha256:22aec2201a63087cc738f7aef642fef9cab957fecfebd58f3019d9f88ef07c1c

Observation 25d3bf16-cb26-4bc2-824e-d762ec0dab1e · outbound

This paper cites Data Filtering Networks.

Multimodal semantic retrieval for product search Data Filtering Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.211886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.211886Z digest=sha256:2879169d52d8b92cf549782c33d69fc90f0dad1cf3fb9ab561f2c53dbb28ecc7

Observation 1aae9dbf-a023-4c7a-b1e0-c3db88923259 · outbound

This paper cites Learning deep structured semantic models for web search using clickthrough data.

Multimodal semantic retrieval for product search Learning deep structured semantic models for web search using clickthrough data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.486530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.216147Z digest=sha256:5b43a898e1bf10087bb1c014b4a9b9d086cdd52c6e6ef43a56610d4286743e28

Observation a6a95a11-f280-42de-8ed1-84701725b540 · outbound

This paper cites Bert: Pre- training of deep bidirectional transformers for language understanding.

Multimodal semantic retrieval for product search Bert: Pre- training of deep bidirectional transformers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.473588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.220141Z digest=sha256:7e0fffd9af162c9d3af76c9072f4c6a96cf81fd8e8090495b3deb364c20dccb5

Observation 083b94a2-312b-4c4d-808e-d54034a2d0df · outbound

This paper cites Relevance-guided supervi- sion for openqa with colbert.

Multimodal semantic retrieval for product search Relevance-guided supervi- sion for openqa with colbert

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.461330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.223612Z digest=sha256:860e7cf5529e6a34f402ffa40c54517c7ece8ad8e54d5afdb18ae935cad13bc1

Observation f25b13c2-3197-4bb3-819d-70570b8545df · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

Multimodal semantic retrieval for product search Vilt: Vision-and-language transformer without convolution or region supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.227144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.227144Z digest=sha256:3503ae14a979324001970a91045412a61b4e0fd378d0989cb74d0045a6429186

Observation ebe36528-50ff-495e-94c7-8dbab2fbd069 · outbound

This paper cites Embracing Structure in Data for Billion-Scale Semantic Product Search.

Multimodal semantic retrieval for product search Embracing Structure in Data for Billion-Scale Semantic Product Search

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.230595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.230595Z digest=sha256:2ad8ddb805dbd79f9bf552dbb85d80aebf852d641241acb9a2848008e3956df5

Observation 86850806-4794-49fc-9955-0b4b98ea031d · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and gen- eration.

Multimodal semantic retrieval for product search BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and gen- eration

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.441106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.234744Z digest=sha256:6d2f7d6fa26746b262ba9705f5185a992d0a3d50e4914df4fa04bf34e4b1d71d

Observation 8001befc-39be-49df-b6c9-b9a088a1e7af · outbound

This paper cites Embedding-based product retrieval in taobao search.

Multimodal semantic retrieval for product search Embedding-based product retrieval in taobao search

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.430600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.238434Z digest=sha256:2c577419754de78ad5f892e7eca2f01a23da95060949b2647ace49bfc4684125

Observation a2f4c79a-eda3-4a11-8c59-2521b90486fd · outbound

This paper cites Deep self-adaptive hashing for image retrieval.

Multimodal semantic retrieval for product search Deep self-adaptive hashing for image retrieval

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.419771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.242178Z digest=sha256:8a2f8e72559235a32808b32fb0ec711f98b7c1e496f38612edde09f45a79fcea

Observation 02e29bcf-64c3-4a95-ab97-dd2103fbad97 · outbound

This paper cites One picture is worth a thousand words? the pricing power of images in e-commerce.

Multimodal semantic retrieval for product search One picture is worth a thousand words? the pricing power of images in e-commerce

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.407681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.245939Z digest=sha256:dd4b281d8c9c492a50a07a10c5741ea3e16dc0d9c2e7db388a24a43c40049c79

Observation 79336b39-1048-4115-be30-0bbf3813f663 · outbound

This paper cites Dc-bert: Decoupling question and document for efficient contextual encoding.

Multimodal semantic retrieval for product search Dc-bert: Decoupling question and document for efficient contextual encoding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.395179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.249501Z digest=sha256:c6332394150afc844f61f24dd4da9601700c38e28ca8a1fbb2bd27198b62e7b9

Observation 0bc18342-8fe9-4e35-abde-c93e351e474a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Multimodal semantic retrieval for product search Learning transferable visual models from natural language supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.253172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.253172Z digest=sha256:7708c0445f27f2b54acd070d6813be6885f85612c0f18be0ac13217784258598

Observation 17e39bd2-40ae-4e50-96fc-609d8e7e6dc5 · outbound

This paper cites Introduction to information retrieval, volume 39.

Multimodal semantic retrieval for product search Introduction to information retrieval, volume 39

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.256947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.256947Z digest=sha256:81194b8176e5b69ce92dee41adb6a0caad84126f7a90092260cbf3e01a155f00

Observation 0047cd88-5228-4bd7-83ab-c9c093c0b4d8 · outbound

This paper cites RepBERT: Contextualized Text Embeddings for First-Stage Retrieval.

Multimodal semantic retrieval for product search RepBERT: Contextualized Text Embeddings for First-Stage Retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:16.260174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:16.260174Z digest=sha256:2ed83f97752dbcd751fd9ca8f2b3b87ed8d3b6f35165ccf608436efd547589f0

Observation fcce507d-7469-486a-99cc-532d1796031f · outbound

This paper cites Binary neural network hashing for image retrieval.

Multimodal semantic retrieval for product search Binary neural network hashing for image retrieval

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.368497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.263965Z digest=sha256:0c3f0d6fee1991444e9e055a6b8cdcaadaf59ab1292159c36d15c0e03a6e197b

Observation 3ae9d517-66e5-4be6-93ed-e85810624113 · outbound

This paper cites Bringing multimodality to amazon visual search system.

Multimodal semantic retrieval for product search Bringing multimodality to amazon visual search system

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.357561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.268023Z digest=sha256:363f33de89900ac2a5e382d9c75e1d534651de4886543d3e661a508b469cd8e7

Observation cb9e5e22-5576-40eb-9b4d-dc6446244ec3 · outbound

This paper cites Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks.

Multimodal semantic retrieval for product search Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:46:16.346018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T20:46:16.272642Z digest=sha256:a0dded0c5b24bda134734ee822d56821f107d555e0c33fabca01d5a6be3acf26

Pith citing papers

No inbound Pith citation observations are available.