Pith. sign in

Paper Citation Record · LEDGER

On the rankability of visual embeddings

As of 7 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2507.03683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03683 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:11:51.717772Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact11
  • verified fuzzy29
  • unresolved29
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72fcb7d1-1878-4a5b-987a-7a8956bbd1b7 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

On the rankability of visual embeddings Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:48.623426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:48.623426Z digest=sha256:267ff46aa42a6b862ca269f4bfa96b58b8d187aa70bc14c3f24aa96196d68fcc

Observation 72f8616e-862e-4547-92c6-4f8ae4d82f13 · outbound

This paper cites Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features.

On the rankability of visual embeddings Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.679728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.662843Z digest=sha256:2b4298bb5f1cd783f1b56b493ecd67b20b5c538ee87c911c2b18531c00642e76

Observation 671d080d-4f91-48ef-b91b-bc3ff30dd0f1 · outbound

This paper cites Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models.

On the rankability of visual embeddings Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.140579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.742002Z digest=sha256:879c6e45914227b6f6a5d3b4329b2ab8e7b8fb3f68a9a61bac88c35bc330110b

Observation 82b92998-93e2-45a5-9a65-d480c670884b · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

On the rankability of visual embeddings A Simple Framework for Contrastive Learning of Visual Representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.663979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.803202Z digest=sha256:27c13d8bea2a22bc96d1436dc62624b911bf183db850184dc9014a04488a8c4f

Observation 511cf3cc-d07f-4c12-a0ae-b1d11993093b · outbound

This paper cites Deep Learning for Instance Retrieval: A Survey.

On the rankability of visual embeddings Deep Learning for Instance Retrieval: A Survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.647220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.894792Z digest=sha256:f50bb78c9f9b45bc8b467319a1cfe6a8d02315f0bc4c8dad038db3f10d859d9e

Observation 0208f5c9-189e-4893-a53b-1e5093e7b4a4 · outbound

This paper cites Deep learning for instance retrieval: A survey.

On the rankability of visual embeddings Deep learning for instance retrieval: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.630656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.981497Z digest=sha256:373d15eaa0be004a028292c47497ab45de8948132aecf0b23c50a32828d60cba

Observation af97d169-f7c6-4cc0-8ed5-efa72c98a39c · outbound

This paper cites Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds.

On the rankability of visual embeddings Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:49.072090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.072090Z digest=sha256:47880e97175d49f99b92ebd2aedab8e9a514bb61b1697cc4a7ef287127de1092

Observation 94a871ac-e6e3-41c2-b3b3-06ddd5a78bde · outbound

This paper cites Hyperbolic Image-Text Representations.

On the rankability of visual embeddings Hyperbolic Image-Text Representations

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.112878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.128084Z digest=sha256:fc97ae4494b26943374bbbeb2322194dffd54fac8b129efee1e3eb7377c2843e

Observation 2d825a8c-c3ea-4a01-a78d-877669bf5ce3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

On the rankability of visual embeddings An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.229769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.229769Z digest=sha256:bab6b98643c1236ba56d68e4da2f0b1be2804658fc6ef50d90392861f7d252af

Observation b6cc826c-934f-489b-94e1-f44b9422ddc1 · outbound

This paper cites Teach CLIP to Develop a Number Sense for Ordinal Regression.

On the rankability of visual embeddings Teach CLIP to Develop a Number Sense for Ordinal Regression

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.059438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.329152Z digest=sha256:42852012dc70677fc31125f44f732580bd3a715dccdc50e31061290fc33c50a3

Observation 8883acf4-b68b-4343-903e-8706086d6021 · outbound

This paper cites Age and Gender Estimation of Unfiltered Faces.

On the rankability of visual embeddings Age and Gender Estimation of Unfiltered Faces

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:53.088420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.450137Z digest=sha256:e85efb47c990483768164e399d081d949371d31ff3aa917ff37aa173bc1e530b

Observation 9cfaf071-0d3c-4582-83b4-b5dd8d280749 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

On the rankability of visual embeddings It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.531103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.531103Z digest=sha256:381dad483a37def4722d641b414858434d933ff2517ef79437691aabc4e0a462

Observation 9dd7610a-3d37-4134-a268-e1343c54c0f3 · outbound

This paper cites Heterogeneous face attribute estimation: A deep multi-task learning approach.

On the rankability of visual embeddings Heterogeneous face attribute estimation: A deep multi-task learning approach

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.613502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.721697Z digest=sha256:46d634e622e3cc01bdc3b73034102024614d1ee01afe3f17a1c515fa880dc0c2

Observation 477da1c2-50ff-4f16-8b1d-05c3201567d5 · outbound

This paper cites Deep Residual Learning for Image Recognition.

On the rankability of visual embeddings Deep Residual Learning for Image Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.805256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.805256Z digest=sha256:5e5f5ff8d07402d4677347232a080ad5a849e8c31780f9cb0a3ba0e536e170f0

Observation 307ea798-7ea6-4281-aa6e-2432bac0103f · outbound

This paper cites Multilayer feedforward networks are universal approximators.

On the rankability of visual embeddings Multilayer feedforward networks are universal approximators

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.597532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.928316Z digest=sha256:87b4a5d784dff8f8bd3a1356990b2d74249ea54be8471b851bbf6622ffb6de67

Observation bb2b5fda-2abe-47f1-b5e9-2c4c8f7635f0 · outbound

This paper cites KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment.

On the rankability of visual embeddings KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.582027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.055806Z digest=sha256:488b0b5cc8bbaa229d5de32652b3de8e9c2935a7521f07008d94efc7b4a72c3c

Observation 97fb22fc-96d4-4760-9b18-e002309f4483 · outbound

This paper cites Lp++: A surprisingly strong linear probe for few-shot clip.

On the rankability of visual embeddings Lp++: A surprisingly strong linear probe for few-shot clip

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.566871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.309455Z digest=sha256:f1e40b084fef1e565e8f1a66e465a25a1cb362ca55757bf83f686e85b7755ea4

Observation 9d9125c5-7393-4fdd-90fd-9213f230accf · outbound

This paper cites The Platonic Representation Hypothesis.

On the rankability of visual embeddings The Platonic Representation Hypothesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.427501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.427501Z digest=sha256:006a7190a1644935d32196596989883cf6c33c704e734c04dfc61e935a99ac0b

Observation 1e02910c-73b7-4adf-a889-8aa29d70c8f0 · outbound

This paper cites CLIP-Count: Towards Text-Guided Zero- Shot Object Counting.

On the rankability of visual embeddings CLIP-Count: Towards Text-Guided Zero- Shot Object Counting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.544893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.544893Z digest=sha256:b6e34dcda5e02a4166a4e293e7ea7a1065158ceb215cc3620df69f7aeebf971d

Observation 22adfc6d-f764-4110-8eea-9146e82b38b7 · outbound

This paper cites Hyperbolic Image Embeddings.

On the rankability of visual embeddings Hyperbolic Image Embeddings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.549000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.625607Z digest=sha256:e68090936b99a72c4cdad29165359442f1dd9ac1f83bd6b109befeea82f5fa02

Observation 86d92a71-4aea-4f83-ad09-41ba0cd792a0 · outbound

This paper cites MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs.

On the rankability of visual embeddings MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.693133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.693133Z digest=sha256:15135793bd673765a17c09a8fba38d22d9cda59048cdbdcad0e41ff1b9f1b981

Observation 4c0d978c-48c2-461d-9110-e11b77b6153b · outbound

This paper cites Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav).

On the rankability of visual embeddings Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.529295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.813651Z digest=sha256:b9b89d9e9771953e9cc69fad50b2e920898abc411426204ec333cbbad6e23f2c

Observation d74f666a-fe1f-4221-be11-5cb58b0a6fd4 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.960886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.960886Z digest=sha256:5d277e02f33dda3a63cda5e44c06a1035fd532d392a6fcfc7503e7840260fa9e

Observation de4d0571-cc07-4331-a0c2-034a6b53f7b1 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.002032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.002032Z digest=sha256:81edc2c43796688fa580c7d6abfe65af60573d8bae6ca971c6db5eb571494348

Observation d160e3ef-edc5-4886-b0e0-5c3704ab4ed2 · outbound

This paper cites Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning.

On the rankability of visual embeddings Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.512741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.033881Z digest=sha256:401e33241163e14f3d030b49f7f6d9ae119942820115c01fa9071312e5a4c6b5

Observation c5ada6c4-84e0-41ad-b510-1f10d8b09f98 · outbound

This paper cites MiVOLO: Multi-input Transformer for Age and Gender Estimation.

On the rankability of visual embeddings MiVOLO: Multi-input Transformer for Age and Gender Estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.039704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.039704Z digest=sha256:44bc3240fc8b2e02e1131f1e1aec800600583fd2f02e5e54d5c5ecf84b1cbad6

Observation 9e393176-8412-4a5a-816e-ea283721eaac · outbound

This paper cites The Double-Ellipsoid Geometry of CLIP.

On the rankability of visual embeddings The Double-Ellipsoid Geometry of CLIP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.045027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.045027Z digest=sha256:74b79d96db9dc172eba06a8b8e93ab5292d3f03a56ffbc0c36fe268fc8e44e1e

Observation 4d502446-6440-47de-b3a7-37f7735ac5f9 · outbound

This paper cites Does CLIP Bind Concepts? Probing Compositionality in Large Image Models.

On the rankability of visual embeddings Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.071811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.071811Z digest=sha256:abad919c803a52f496cca8462db8d26f7144c4a9f9c57ec6a17dfb0b97681b3b

Observation 2624fd97-d6cf-4533-b4f8-02e1e764d3dd · outbound

This paper cites Align before fuse: Vision and language representation learning with mo- mentum distillation.

On the rankability of visual embeddings Align before fuse: Vision and language representation learning with mo- mentum distillation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.494930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.171974Z digest=sha256:1c058566324acc51e35b58b52a2256652de8a847deab584b1643ee34670911e7

Observation bf413c30-217d-4123-95e4-6fe4efb58ffe · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation.

On the rankability of visual embeddings BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.478419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.232772Z digest=sha256:1bf5b6adfea926cc4b51b4cd0d67c89ba8d84b98d7014b90d1d67113039167f1

Observation f366b8f4-65fd-4e5b-8ae5-52c09621ca3b · outbound

This paper cites OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression.

On the rankability of visual embeddings OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.956760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.302801Z digest=sha256:a7ec7a60ef45736299bb1888b5232d49cc3dece21648e6e926b45507f159d99b

Observation 00efac1b-a70d-4467-b8f4-0ba5c81b0c6d · outbound

This paper cites CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model.

On the rankability of visual embeddings CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.933505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.386439Z digest=sha256:2891fd18ad7b230283de9dbeeefc70088d63cb1b91a04223da9b830310f6ee76

Observation ffb1b424-5803-447c-983e-274b6ec42009 · outbound

This paper cites Beyond comparing image pairs: Setwise active learning for relative attributes.

On the rankability of visual embeddings Beyond comparing image pairs: Setwise active learning for relative attributes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.462867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.467629Z digest=sha256:c6ecb05823ac7badb061843a5f14fef912a57a76c155a6edf1c77ea6cd6688fd

Observation 9cdd71ea-a5ce-4033-93fd-6d02518fd554 · outbound

This paper cites CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification.

On the rankability of visual embeddings CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.500131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.500131Z digest=sha256:b4613db87cbe2d3304e1025d666ce0573410fe5059fef1a0e96c122d8490b25e

Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.511161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.511161Z digest=sha256:b8cda123b6b01ca83010d8965c2a34dbd97e8b76a293d0fb0010a879918f1e58

Observation ead6e43e-4d76-4484-8b92-a25d7b2ded35 · outbound

This paper cites A V A: A Large-Scale Database for Aesthetic Visual Analysis.

On the rankability of visual embeddings A V A: A Large-Scale Database for Aesthetic Visual Analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.516129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.516129Z digest=sha256:b19a0491c1601a2892831cfd15f4b89d5a17809b4237778652c7da4de68ae4ef

Observation 55c49fff-8449-4320-8028-83eba01abd3f · outbound

This paper cites A metric learning reality check.

On the rankability of visual embeddings A metric learning reality check

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.441046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.520788Z digest=sha256:f760acfaaa52fd0dba193b214e09ad3f50b0ba0d42c244aba7a22a17b87b3741

Observation f8ee04c5-2aa8-46e1-be7d-9572b7fa1707 · outbound

This paper cites Parts of Speech-Grounded Subspaces in Vision-Language Models.

On the rankability of visual embeddings Parts of Speech-Grounded Subspaces in Vision-Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.533367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.526319Z digest=sha256:898a35ea0444093ed0d01bc741ab65f72a8965c514f2143536c1e9bf0c0b65b1

Observation 18e7cfd6-4087-4ef9-beec-b45a1a211e6a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

On the rankability of visual embeddings DINOv2: Learning Robust Visual Features without Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.531605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.531605Z digest=sha256:2783aa8e4f1b3ee67611e3f47ffc7449efba7e87c6137c80276d4d30dfad7f69

Observation bd3a1569-da58-45a3-826f-5ab8b6110eaf · outbound

This paper cites Teaching CLIP to Count to Ten.

On the rankability of visual embeddings Teaching CLIP to Count to Ten

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.537121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.537121Z digest=sha256:c1c039407a5cce984f7010fa06a3b4655e4ca0ea3d7b5ae10b9ce42c0af39f04

Observation 75fa21f4-3934-4458-8b3d-b3fd6aaebe20 · outbound

This paper cites Dating Historical Color Images.

On the rankability of visual embeddings Dating Historical Color Images

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.542600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.542600Z digest=sha256:04fe6be5db14890f7be41f0d6faabbea6c0349b0fc99d88da22c4cf92a4d646d

Observation 5e8fa665-8c00-4a0b-b22f-f10365a35b98 · outbound

This paper cites Relative Attributes.

On the rankability of visual embeddings Relative Attributes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.548487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.548487Z digest=sha256:0cf1548c4f329d26e5cc5b434fe78426a17487ca327dbff031277367a41329c7

Observation 0babb719-b950-468c-a839-5b56f2279032 · outbound

This paper cites HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports.

On the rankability of visual embeddings HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.422955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.553683Z digest=sha256:76087880aa6beccfbc1ebe2a048c6d9af35f69b8c6563d17d640cb12e573aaca

Observation a0e88aef-b4aa-4aff-90ec-4a7203a7dd5e · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.559058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.559058Z digest=sha256:98e366c47cd34e81205bb6a7e086cf3c9167f7f954675ff8e1e3c1d04237f6f2

Observation 025ec07b-d7e3-459d-99a9-9abf9a1a44e9 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Super- vision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.404122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.564088Z digest=sha256:b2517046dc1e378048e701156c3a23aafed44cd73ceb736137623efcac2661cc

Observation e2d4c7c3-cb2a-4eed-bf32-218603569642 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

On the rankability of visual embeddings Steering Llama 2 via Contrastive Activation Addition

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.569573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.569573Z digest=sha256:8cca46c4464d4fda713d88502af8e50a078ef80ebc3f8ec19e10c9cc5235f9a9

Observation fe3b8e31-c769-4e73-ae32-dd0e949665a4 · outbound

This paper cites Finetuning CLIP to Reason about Pairwise Differences.

On the rankability of visual embeddings Finetuning CLIP to Reason about Pairwise Differences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.575686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.575686Z digest=sha256:66792bb8b52cf5bf3187fc1b51a41c9badb678fe9ddb7b60b072b492500b13d8

Observation 3fd03dbe-f488-4d9f-be8d-8c43e02a1b47 · outbound

This paper cites Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification.

On the rankability of visual embeddings Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification

Reference 49

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T20:11:52.391746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.583085Z digest=sha256:61d107122d0c1aaffebcf67663d163c5a3ebdee95832d49154bbb2cfc748e6a7

Observation bae4703a-499c-46ee-8035-7a7ab5c85cd7 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.386329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.588973Z digest=sha256:e265c7b262ae9f548a0729b196b070b554a1e99fd771f233e3aa9dc89273af77

Observation 3238f178-81da-4011-99c2-d6fe4be79ac2 · outbound

This paper cites Linear Spaces of Meanings: Compositional Structures in Vision- Language Models.

On the rankability of visual embeddings Linear Spaces of Meanings: Compositional Structures in Vision- Language Models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.368075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.601038Z digest=sha256:f666aaafcd57ae5166a486921940a210ff699ceebfcaaf8cd7ee5370f219c79d

Observation 1b0cdd12-01f9-4d10-9361-19f08ebb1ee4 · outbound

This paper cites Deep image prior.

On the rankability of visual embeddings Deep image prior

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.348792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.607467Z digest=sha256:b72eeb971a25fa7d4203a2e792603ea3203a5f3039108eadd0a845aceb7b52a3

Observation e4fe7ec6-2185-47a2-af26-0d7a5a3c787e · outbound

This paper cites BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION.

On the rankability of visual embeddings BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.330263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.612865Z digest=sha256:eff23e1183fe3ea76a03a6ec470bf715dd93c2e7c3e3932402adb44cc7b237d2

Observation 4230b46f-7018-4a08-8e07-12086fe5af61 · outbound

This paper cites Intermediate Layer Classifiers for OOD Generalization.

On the rankability of visual embeddings Intermediate Layer Classifiers for OOD Generalization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.308628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.617981Z digest=sha256:37b954cb77bcca96154964771e631863066d7d8a96eba8b14d623dc7caeed9a6

Observation 9247635d-3ab4-46b7-847f-bf53ff49548e · outbound

This paper cites Order-Embeddings of Images and Language.

On the rankability of visual embeddings Order-Embeddings of Images and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.623271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.623271Z digest=sha256:92ad8214110a39fc2776daf0a569aa13f70772b4a203040d1b057550f34d424d

Observation 6b5e8000-421a-459c-a408-8b76f589aed6 · outbound

This paper cites Exploring CLIP for Assessing the Look and Feel of Images.

On the rankability of visual embeddings Exploring CLIP for Assessing the Look and Feel of Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.628732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.628732Z digest=sha256:dce328e0c137aad214e21ffecab0bec7ff857e98741d6912696455351555f816

Observation 936ac420-362e-4fbf-8e94-e91a8469f375 · outbound

This paper cites Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification.

On the rankability of visual embeddings Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.290182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.634171Z digest=sha256:bbc44a6c13a16496074e2c6b4fe8b43de05ad4b2e1a6de0ef5097b84d6b2d056

Observation e4732314-04f8-4044-ba8f-e92acfb33c58 · outbound

This paper cites Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification.

On the rankability of visual embeddings Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.271360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.640208Z digest=sha256:7a3f7c29cab6a2f1875feb5aede73a31fbd9be96ffde6fd805171f9089ff08d4

Observation be34882b-7800-4c43-bc56-e814d8e5511c · outbound

This paper cites Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere.

On the rankability of visual embeddings Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.646034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.646034Z digest=sha256:629d89cdf5175541754ecde70edb062f129e838642218e9520f313d896d8ac24

Observation 6f1bb15f-0122-4804-9671-9633a12c7e06 · outbound

This paper cites Disentangled representation learning.

On the rankability of visual embeddings Disentangled representation learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.252755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.651508Z digest=sha256:69b03b07aa0aacd34eedd1b7d91240e8c1f500ed5d58e1333cba349fc9cb17d3

Observation f0ff0cf9-3c50-4858-9037-a99cb6c6d5d9 · outbound

This paper cites ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders.

On the rankability of visual embeddings ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.656002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.656002Z digest=sha256:12bae7431bd38ddbf312338b4caa826d30e7d2e5344b8a43dfd565ee77ab4008

Observation c36c12cd-79c4-42e7-a276-151ea491f4ee · outbound

This paper cites CLIP Brings Better Features to Visual Aesthetics Learners.

On the rankability of visual embeddings CLIP Brings Better Features to Visual Aesthetics Learners

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.250347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.661258Z digest=sha256:830171f918ad1f2d26d81c6caf13f5e476546ffd5664d133f8f8c1380f4ccd05

Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.667082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.667082Z digest=sha256:e33527345f0f2682838014ab6b5aa3b0bc8f2154e558b80800610157b06ea2ea

Observation 01c94fe3-87ce-4cef-8ba2-51249241ef8a · outbound

This paper cites Just noticeable differences in visual attributes.

On the rankability of visual embeddings Just noticeable differences in visual attributes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.233523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.673658Z digest=sha256:c38fa93a9daac3bd9ee36b7c550bb742b5ce86043ec7610f8b6c564cfe168270

Observation 82ad9c1e-42a9-416c-b056-06aac0007fb0 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

On the rankability of visual embeddings CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.216137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.679214Z digest=sha256:db9cb7c7f9183518a8a6e307e2368645f2683cdfe054c897006dc258026ef9fd

Observation db9bfc04-6741-4d31-9d31-7c52aab8b687 · outbound

This paper cites RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP.

On the rankability of visual embeddings RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.200041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.684176Z digest=sha256:6972281a0cd6a56bb90cc2db534914fef90c8a5ce273ab05e0442bbc033dbea9

Observation 3f14a96a-9b7b-416c-97c0-65bfe82f4afc · outbound

This paper cites Ranking-aware adapter for text-driven image ordering with CLIP.

On the rankability of visual embeddings Ranking-aware adapter for text-driven image ordering with CLIP

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.801956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.689502Z digest=sha256:9fccc464c678a78ac26c5dea2a18f04624e39884b6d4d1ea0e2dbe05dca45978

Observation 3ed8fddf-ba6d-48cb-8c12-f3952ea698e1 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

On the rankability of visual embeddings When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.695389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.695389Z digest=sha256:a73bf979b6e35e2ac9fe95c8747bd5727ee61592aeec9cc498145bc8e1d737de

Observation 68d75df2-e218-4e43-b214-cd4e2f6db13d · outbound

This paper cites Single-Image Crowd Counting via Multi-Column Convolutional Neural Network.

On the rankability of visual embeddings Single-Image Crowd Counting via Multi-Column Convolutional Neural Network

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.701459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.701459Z digest=sha256:b218ad4b63540b4a285c372c95d45ed29ecd9943e145a8aeebea893e7382804d

Observation ef6fcadf-a217-4378-ae55-e1e091615786 · outbound

This paper cites an unresolved cited work.

On the rankability of visual embeddings Unresolved cited work

Reference 70

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:52.187722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.706634Z digest=sha256:a31fedf0c450a9739e312307da2b7e9efb991501798108e8d3797630fa80469b

Observation e330120c-f2ee-478c-809d-c67b5537b483 · outbound

This paper cites Learning Ordinal Relationships for Mid-Level Vision.

On the rankability of visual embeddings Learning Ordinal Relationships for Mid-Level Vision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.183574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.717772Z digest=sha256:0ed85e1f51d1d4eddbc1dbf07d03b36e2505db897bff56b774fe1e9496a4dd9b

Observation 73c93712-e3f5-4205-97ab-3869d8d60f4d · outbound

This paper cites Age Progression/Regression by Conditional Adversarial Autoencoder.

On the rankability of visual embeddings Age Progression/Regression by Conditional Adversarial Autoencoder

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.767240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.712243Z digest=sha256:44530b46ce25ea2e3b8f135ab3bb1da046dbad836da4592860b67e7bfa81c724

Observation ec1d0391-a4a0-4c85-84d5-c2f4ae582f49 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.594734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.594734Z digest=sha256:546b521b2a9b7d4a85b7b0d50829cc510920798fcb480ce253689479ae489491

Observation 7d03c2f3-6c39-41dc-89c0-7120d1ade950 · outbound

This paper cites DOI: 10.1109/TIP.2020.2967829.

On the rankability of visual embeddings DOI: 10.1109/TIP.2020.2967829

Reference 4056

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.176194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.176194Z digest=sha256:85de614ebc4e33a21ef7bb5371040ef92b74102c3c02055f03e48ad95d03b8f9

Pith citing papers

No inbound Pith citation observations are available.