Pith. sign in

Paper Citation Record · LEDGER

Seeing the Abstract: Translating the Abstract Language for Vision Language Models

As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2505.03242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03242 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:02:33.204350Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c3c98f7-a1d2-4e9d-a707-2f637b9a84b2 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.011938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.011938Z digest=sha256:dd8a2b0628581be6bfdc78444383b25890fc4731b8f2311b0d3f3388910e8fa5

Observation 1130a012-6b04-4a0e-b790-8869522511f6 · outbound

This paper cites Fashion product images (small), 2019.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Fashion product images (small), 2019

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.957305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.018210Z digest=sha256:39f9da9e224080ddd494143c29a56d283378f2cf1859c4559df61a4c1f46351b

Observation 50bee6d9-7a3c-41ab-aca5-b369a354f40d · outbound

This paper cites Compositional learning of image-text query for image retrieval.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Compositional learning of image-text query for image retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.943015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.022904Z digest=sha256:f70440c21ea6ffb85ee6e33d4731940e6445740bbd4fa9beb62ed00afe8110ce

Observation 8aee5ede-c3aa-4363-8732-421ef2f74c86 · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Effective conditioned and composed im- age retrieval combining clip-based features

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.928127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.027612Z digest=sha256:01e9f62a2e4d4eb1a4441bcb8dfa8754dc311878a40a8d3010a0aaf25a5cfe2c

Observation 0d88aab7-7541-46e7-b399-203b1865c234 · outbound

This paper cites Concreteness ratings for 40 thousand generally known en- glish word lemmas.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Concreteness ratings for 40 thousand generally known en- glish word lemmas

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.913521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.032117Z digest=sha256:7f3eea6bab351a7617f9dfffed63a7bc7a9b03dede1160597eb4ffbfe7c251df

Observation 5fba8645-11ae-4967-a6cc-6a57971e6d0a · outbound

This paper cites Openfash- ionclip: Vision-and-language contrastive learning with open- source fashion data.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Openfash- ionclip: Vision-and-language contrastive learning with open- source fashion data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.898808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.036422Z digest=sha256:eb748a3c556db096b302141dba887c83ccc11562205d481cce9ddba344bbd722

Observation 337e550e-c887-4f49-9a42-8654823ef4f8 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Reproducible scal- ing laws for contrastive language-image learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.884496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.041454Z digest=sha256:fcaf4f1be7599b5cf5b52d053be55108f8c50c8a282cb5e1113f2750093c4192

Observation 0431e9b4-0571-40da-85bc-7d18572a29b2 · outbound

This paper cites Contrastive language and vi- sion learning of general fashion concepts.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Contrastive language and vi- sion learning of general fashion concepts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.869532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.045877Z digest=sha256:69ee1149cdadbe764a1b4fefe8be15fe0a691bc0398fb8124b3dfb542c2bb915

Observation 2faacbe2-f26e-446d-8e52-2b37b459ee0e · outbound

This paper cites Style finder: Fine-grained clothing style detection and retrieval.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Style finder: Fine-grained clothing style detection and retrieval

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.855069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.050148Z digest=sha256:37302db5b5d9558138f86bdd1d21e7abc5bba10165e60191bf91354f05dcf643

Observation d0d34711-bba8-4f6a-87a9-ae0859429f8c · outbound

This paper cites A Survey on In-context Learning.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models A Survey on In-context Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.054416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.054416Z digest=sha256:ac02b7ec79e6d70456c67d3edb05f4b6936fc80f702da3316917da4d7740cc5d

Observation c3249b55-16dd-4104-939a-830e141103be · outbound

This paper cites The Llama 3 Herd of Models.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.058679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.058679Z digest=sha256:afadb3a66dba7df6a9c51fbe9ffbb4af3654230531d122440277bf1543decd3e

Observation 0b498e6e-1c3d-4c47-91a0-cd115027dfc9 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Dat- acomp: In search of the next generation of multimodal datasets

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.840254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.063314Z digest=sha256:c13c4d2c335c8847df66ab2c5642ff664eec23c38aeb207e9362d874c37ae9d2

Observation f600fd3a-d9a6-45a9-8c34-bcd5a3a5469e · outbound

This paper cites Fashionvlp: Vision language transformer for fashion re- trieval with feedback.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Fashionvlp: Vision language transformer for fashion re- trieval with feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.824795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.067832Z digest=sha256:d4eddc433d2a5d7f5e10e5311aea202933ac698032b7db52db6d2cc6809a7c2e

Observation 249663be-8ec0-4874-9c23-d5aa3a20790d · outbound

This paper cites Scott, and Serge Belongie.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Scott, and Serge Belongie

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.809323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.072286Z digest=sha256:1a5f0a4bb5bf141334ded30c68e3a836570dfd4fc93302340afbb37088da0ef6

Observation df531b13-ec87-412a-b06b-0609cbe5beac · outbound

This paper cites Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Huang, Xiao Zhang, Menglong Zhu, Yuan Li, Yang Zhao, and Larry S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.793548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.076495Z digest=sha256:1df7dc48734c1e6c89d4ec5b8ae4182cba798aa22b2b027724eda034930e1861

Observation 9df7489e-d73a-47e5-b022-9302eb3d497e · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.080711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.080711Z digest=sha256:32bd57b683b33bb04d969d6ed719b0f228918467457c64e62b35277227e2c117

Observation 92d27575-e7fc-4fae-8d1d-3faab532a7e0 · outbound

This paper cites spacy: Industrial-strength natural lan- guage processing in python, 2020.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models spacy: Industrial-strength natural lan- guage processing in python, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.779144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.086130Z digest=sha256:b296cf1e81aece63d8d82676ad02939e42186b4fc548758afe643fdac843fbd7

Observation f95ba6e1-83c8-4947-9dc3-5da770efa273 · outbound

This paper cites Feris, Qiang Chen, and Shuicheng Yan.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Feris, Qiang Chen, and Shuicheng Yan

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.764540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.090493Z digest=sha256:892417a640c44bb2879ad9c720b65ab0fc9b89d1454b76a8d856d93fb27d1fb4

Observation 7416e206-8743-4af8-a246-13c1cffd0a1b · outbound

This paper cites Multi-label fashion image classification with minimal human supervision.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Multi-label fashion image classification with minimal human supervision

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.748411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.095115Z digest=sha256:3f85e05b81b3418d16a7e5e5bfb2e620c70f8c0ceb70600acdacb7704d140547

Observation 4ee5b739-5acf-41c6-8172-e5e8ffce4e64 · outbound

This paper cites Cross- domain image retrieval with attention modeling.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Cross- domain image retrieval with attention modeling

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.733277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.099520Z digest=sha256:42b448cd79f6ffb0123ed3d1bd11027f5e7f06a9ffcdac1e6e1d344d12cf100c

Observation f32aadfb-536c-4cdb-a4d0-c6fda574d590 · outbound

This paper cites Hadi Kiapour, Xufeng Han, Svetlana Lazebnik, Alexan- der C.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Hadi Kiapour, Xufeng Han, Svetlana Lazebnik, Alexan- der C

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.717084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.103752Z digest=sha256:32bc75432e7a8a25fb3ec908b61dadcd1a0a0a8ec1ca3f7f1da9a065a8c0ae61

Observation a2bad5db-f6a6-4c8e-a584-3b9482b58c6b · outbound

This paper cites Cosmo: Content-style modulation for image retrieval with text feed- back.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Cosmo: Content-style modulation for image retrieval with text feed- back

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.699369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.107981Z digest=sha256:ceb0b3e426b5bdf6d0de744b1d3d5082651345f3192cace6b217bd939af69afd

Observation b4a03a62-29a0-477d-81cf-826b765be2bd · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.680884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.112226Z digest=sha256:1ab78c701159b77ab4aaea4a346781c2b904874cbbea597e4b9783581a088165

Observation fad16657-c19b-440b-b7e9-baa648f0913d · outbound

This paper cites Deepfashion: Powering robust clothes recog- nition and retrieval with rich annotations.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Deepfashion: Powering robust clothes recog- nition and retrieval with rich annotations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.661965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.116734Z digest=sha256:468edd18708b6859ea71000c06cb148565fd2291cac2a645fcdd8078894e2d9d

Observation 15350a28-6f30-478d-9799-db2f8bfe1af5 · outbound

This paper cites Matthews.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Matthews

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.646171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.121328Z digest=sha256:e6998326f0fa1fbf2c50b9e7ae58be97113361a6d8a0393871adea5cbed6b374

Observation fc282812-771b-4be0-933b-93527640cd31 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Learning transferable visual models from natural language supervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.630334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.125608Z digest=sha256:acc307851adccc1abb4c6ffd69b1d3630631b30d3e7bf66cdb9c5e8cc5fc4795

Observation 77883fbb-2104-4932-9fc4-3d90c683db2c · outbound

This paper cites Fashion-Gen: The Generative Fashion Dataset and Challenge.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Fashion-Gen: The Generative Fashion Dataset and Challenge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.130036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.130036Z digest=sha256:5917bc4dbc6444ea5fc4c1a52d4db1557d9f03d3b090bef7d7679e9fbefa0fc1

Observation 2609d7e3-60ec-46c3-ab51-61d4240b510d · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.136073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.136073Z digest=sha256:ce15b3f86b29eb750fda04b30c9e3d3cffceb545a1670db25bda5900bacec2f3

Observation 4cb6e6fd-02ba-495e-af9a-cf708361f448 · outbound

This paper cites A Tutorial on Principal Component Analysis.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models A Tutorial on Principal Component Analysis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.140679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.140679Z digest=sha256:fd687a7154515869722f9f4462d4af34cafd3de2eab8b04c43d106f71123b1cd

Observation eb9a49a2-2496-4219-bc2d-10efc81edfa8 · outbound

This paper cites Neuroaesthetics in fashion: Mod- eling the perception of fashionability.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Neuroaesthetics in fashion: Mod- eling the perception of fashionability

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.613030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.145361Z digest=sha256:ebd5e70d084b63f7e3e0bde3696a0fa292ac15769385fce44710affefcd826ab

Observation 36c1c18d-518c-4834-be3d-a8254071c761 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.149592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.149592Z digest=sha256:26b3cda485805d23c49f4d25b104409a3cc904da667c7bb035831011664dcacf

Observation ee2e54c3-10e0-4dbe-9ba2-a73ff3eb8300 · outbound

This paper cites What makes a style: Experimental analysis of fashion prediction.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models What makes a style: Experimental analysis of fashion prediction

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.560644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.154179Z digest=sha256:a787eff96aba94a3e9629027b3939f67079399b4a52cafbbb8c425b7fc542089

Observation 6f9c7cd0-b3c0-4d5b-819f-0894c1745cd0 · outbound

This paper cites Visualizing data using t-sne.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Visualizing data using t-sne

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.158480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.158480Z digest=sha256:74dbc9f5727971776ff651ae27ae13c6d90cf32d9452208e606469fa168099bf

Observation b36767c4-1756-487c-89e7-e4c23e966099 · outbound

This paper cites Vinson, Marco Tettamanti, Joseph T.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Vinson, Marco Tettamanti, Joseph T

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.510470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.162757Z digest=sha256:6893715575ce23d994aefe4cf1f37973c79c7627308c3aaf22bc3694e76467e0

Observation 71ca092f-60d1-463c-8b69-cea11389fa4d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.167290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.167290Z digest=sha256:1d45c80890d6bf661c8af86196b126c6e71800b59fbbfc772d2a8f90f0fbae5c

Observation 94290281-25d3-44c4-acf2-f030547900b2 · outbound

This paper cites Clothes search in con- sumer photos via color matching and attribute learning.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Clothes search in con- sumer photos via color matching and attribute learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.494805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.172411Z digest=sha256:6bd8336e80d3ceead347c9663630fec1d184df50ad7339b8804f34ccba036d40

Observation 5c5c8509-ec41-4fe9-bc3d-e749d0159308 · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.478362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.176766Z digest=sha256:94e4572b522ada1fe0e396c0e95e38a43c6366effed0d768347bd13ea43f2283

Observation 9c828452-ad37-4bf8-9d9d-6abac93f0066 · outbound

This paper cites Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:02:33.181886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:02:33.181886Z digest=sha256:f8f835d907d79074782bb4ef5781d79ccf0193a79c82edbc84720a86a4c62bb2

Observation e1350231-81c8-4e53-96f4-11b1d05ff49a · outbound

This paper cites Fashion captioning: Towards generating accurate descrip- tions with semantic rewards.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Fashion captioning: Towards generating accurate descrip- tions with semantic rewards

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.462130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.186359Z digest=sha256:21ef89210fc0600bc986b2a2f855488336990ccda299cb5305c5353f551897ee

Observation a5764240-858d-4821-a85d-cf14fa824389 · outbound

This paper cites Sigmoid loss for language image pre-training.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models Sigmoid loss for language image pre-training

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.446319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.190681Z digest=sha256:2ba3092db046b86e5ae8e5bebfb13a471e2f5e7c36f92037cebc9855c1f6f449

Observation 72bfe066-24e8-43ae-a12e-2ff3e6b3e2ae · outbound

This paper cites chic” and “street ready.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models chic” and “street ready

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.429305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.194952Z digest=sha256:b80cb762e61be635cd5ba2af4d3e59b54b09b694be77bf1a80587884990c0456

Observation 85282655-cebe-439c-93c5-f757ed9367e9 · outbound

This paper cites adjective.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models adjective

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.413730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.199797Z digest=sha256:3407264b3a5f5dced203704598ed6c6d2f6810fcac0c48e86fad3eca1b0cbd1a

Observation 4b95f207-1dc3-45ff-97c3-6179bdfaf0d1 · outbound

This paper cites compound word.

Seeing the Abstract: Translating the Abstract Language for Vision Language Models compound word

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:02:33.396384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T00:02:33.204350Z digest=sha256:e7f8d73257187987420a150499485a187102c7e144c10d05eaa4b738e68839af

Pith citing papers

No inbound Pith citation observations are available.