Pith. sign in

Paper Citation Record · LEDGER

On the rankability of visual embeddings

As of 7 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2507.03683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03683 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:11:51.717772Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact11
  • verified fuzzy29
  • unresolved29
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72fcb7d1-1878-4a5b-987a-7a8956bbd1b7 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

On the rankability of visual embeddings Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:48.623426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:48.623426Z digest=sha256:17cd213fcbb2e93377ac1cdbc9a590fa0464acf8cf0a405586dc43f8aeb1305a

Observation 72f8616e-862e-4547-92c6-4f8ae4d82f13 · outbound

This paper cites Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features.

On the rankability of visual embeddings Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.679728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.662843Z digest=sha256:698e324170d7415753d5ba283aafb10607a85bdc398cc97fd00b39e7bb1af8fb

Observation 671d080d-4f91-48ef-b91b-bc3ff30dd0f1 · outbound

This paper cites Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models.

On the rankability of visual embeddings Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.140579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.742002Z digest=sha256:e28f33527e4b1785df7a81c21a07ae7dbebff240c8115529ed725356270405cf

Observation 82b92998-93e2-45a5-9a65-d480c670884b · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

On the rankability of visual embeddings A Simple Framework for Contrastive Learning of Visual Representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.663979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.803202Z digest=sha256:f2fef7d57c73f78c858ceeb3c0e66bd48b9d09a27018d02eae384631e37c45fe

Observation 511cf3cc-d07f-4c12-a0ae-b1d11993093b · outbound

This paper cites Deep Learning for Instance Retrieval: A Survey.

On the rankability of visual embeddings Deep Learning for Instance Retrieval: A Survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.647220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.894792Z digest=sha256:c66dc50f90dc2ab6f68c4a60063462f0a6682db9a21dbb253809b5f342450b14

Observation 0208f5c9-189e-4893-a53b-1e5093e7b4a4 · outbound

This paper cites Deep learning for instance retrieval: A survey.

On the rankability of visual embeddings Deep learning for instance retrieval: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.630656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:48.981497Z digest=sha256:0650456ab9b4484a1c357d32aebe208c39d127262d3f846b1ee5946bd416d7ed

Observation af97d169-f7c6-4cc0-8ed5-efa72c98a39c · outbound

This paper cites Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds.

On the rankability of visual embeddings Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:49.072090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.072090Z digest=sha256:7dadd19b64c39f3021ae86cb0bbbdbc0035c9e07e982046fd9e4ef38580d332d

Observation 94a871ac-e6e3-41c2-b3b3-06ddd5a78bde · outbound

This paper cites Hyperbolic Image-Text Representations.

On the rankability of visual embeddings Hyperbolic Image-Text Representations

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.112878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.128084Z digest=sha256:59aed3387b5ff17ca4e3f7456355f69ce722c20cb9b2c5ec4f1b260b137c5415

Observation 2d825a8c-c3ea-4a01-a78d-877669bf5ce3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

On the rankability of visual embeddings An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.229769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.229769Z digest=sha256:ec3b09b15440c138c153f853761d2d7ddd0815957d62805ab5ff3fd1c124e9bf

Observation b6cc826c-934f-489b-94e1-f44b9422ddc1 · outbound

This paper cites Teach CLIP to Develop a Number Sense for Ordinal Regression.

On the rankability of visual embeddings Teach CLIP to Develop a Number Sense for Ordinal Regression

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.059438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.329152Z digest=sha256:160ac51cfe2f534550e2806df3bc5c527be7bc6c60bc0cbb81ce41fb04770438

Observation 8883acf4-b68b-4343-903e-8706086d6021 · outbound

This paper cites Age and Gender Estimation of Unfiltered Faces.

On the rankability of visual embeddings Age and Gender Estimation of Unfiltered Faces

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:53.088420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.450137Z digest=sha256:4ba25b30018bd0a8bb3087cc3a49d7367540d5018d8831ca1cbc0aca32477eae

Observation 9cfaf071-0d3c-4582-83b4-b5dd8d280749 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

On the rankability of visual embeddings It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.531103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.531103Z digest=sha256:578a17273c15fd70513c9ec42868b049d9390e5dee42a501764aa9ae58e3a296

Observation 9dd7610a-3d37-4134-a268-e1343c54c0f3 · outbound

This paper cites Heterogeneous face attribute estimation: A deep multi-task learning approach.

On the rankability of visual embeddings Heterogeneous face attribute estimation: A deep multi-task learning approach

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.613502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.721697Z digest=sha256:f98a3a5f7510772dbf93fcb36ff727c8c78cec89fcdc97408edf0ea74c3bbdfa

Observation 477da1c2-50ff-4f16-8b1d-05c3201567d5 · outbound

This paper cites Deep Residual Learning for Image Recognition.

On the rankability of visual embeddings Deep Residual Learning for Image Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.805256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.805256Z digest=sha256:c4d7e1c5a4328062902abf03007b72044a4a7706f34f76c369ab0160b4bb109c

Observation 307ea798-7ea6-4281-aa6e-2432bac0103f · outbound

This paper cites Multilayer feedforward networks are universal approximators.

On the rankability of visual embeddings Multilayer feedforward networks are universal approximators

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.597532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:49.928316Z digest=sha256:275fc3d6462f775c9e6628b8c719c95859d01b7c016d247cf0b25eb205b66463

Observation bb2b5fda-2abe-47f1-b5e9-2c4c8f7635f0 · outbound

This paper cites KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment.

On the rankability of visual embeddings KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.582027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.055806Z digest=sha256:59cdfe8498e46943a02167925115d009a87e1b10870938a364428395785692ee

Observation 97fb22fc-96d4-4760-9b18-e002309f4483 · outbound

This paper cites Lp++: A surprisingly strong linear probe for few-shot clip.

On the rankability of visual embeddings Lp++: A surprisingly strong linear probe for few-shot clip

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.566871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.309455Z digest=sha256:fe1f7ba3bed6b5c1d92ce5623f7d0a93777009e73d326eb230d88d2a1534829c

Observation 9d9125c5-7393-4fdd-90fd-9213f230accf · outbound

This paper cites The Platonic Representation Hypothesis.

On the rankability of visual embeddings The Platonic Representation Hypothesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.427501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.427501Z digest=sha256:fcaf4f23e11cdd44a0ba6f4210580496972510e02edac17722c0300541cc008c

Observation 1e02910c-73b7-4adf-a889-8aa29d70c8f0 · outbound

This paper cites CLIP-Count: Towards Text-Guided Zero- Shot Object Counting.

On the rankability of visual embeddings CLIP-Count: Towards Text-Guided Zero- Shot Object Counting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.544893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.544893Z digest=sha256:683d6a81ecbdb9469cb5bdacff5c9c5fc6e953fdd6943bad80a779f99aee7843

Observation 22adfc6d-f764-4110-8eea-9146e82b38b7 · outbound

This paper cites Hyperbolic Image Embeddings.

On the rankability of visual embeddings Hyperbolic Image Embeddings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.549000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.625607Z digest=sha256:a6c0a9dee7a1cd608c1a700e920254072d2ab0abb9f92b3fa1708e441f877575

Observation 86d92a71-4aea-4f83-ad09-41ba0cd792a0 · outbound

This paper cites MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs.

On the rankability of visual embeddings MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.693133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.693133Z digest=sha256:064390b4a8da17d5cbe77e68af94f2a0a96f21a816234a3777728889c190ec29

Observation 4c0d978c-48c2-461d-9110-e11b77b6153b · outbound

This paper cites Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav).

On the rankability of visual embeddings Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.529295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:50.813651Z digest=sha256:cc40501072cc58d95ba5fa638baa6c063153d6e2721169d2fb69722184466022

Observation d74f666a-fe1f-4221-be11-5cb58b0a6fd4 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.960886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.960886Z digest=sha256:1aed34fedb00720235391dd0a8756fe2b76dd68b9ac34c4c9e093724e048665e

Observation de4d0571-cc07-4331-a0c2-034a6b53f7b1 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.002032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.002032Z digest=sha256:7a106eefc5de0f0e417afcbc6dbed6ad785c16ea8657005a26f53fb3b69fa9e7

Observation d160e3ef-edc5-4886-b0e0-5c3704ab4ed2 · outbound

This paper cites Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning.

On the rankability of visual embeddings Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.512741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.033881Z digest=sha256:ea3884ca2cca1e20415741ab571ba79621807366f8f85dba39c04cb9607fc3f0

Observation c5ada6c4-84e0-41ad-b510-1f10d8b09f98 · outbound

This paper cites MiVOLO: Multi-input Transformer for Age and Gender Estimation.

On the rankability of visual embeddings MiVOLO: Multi-input Transformer for Age and Gender Estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.039704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.039704Z digest=sha256:a4a4181b36e7bb574a7441350c46164c4c85dd526cc5e55e2379e1b7e1185c23

Observation 9e393176-8412-4a5a-816e-ea283721eaac · outbound

This paper cites The Double-Ellipsoid Geometry of CLIP.

On the rankability of visual embeddings The Double-Ellipsoid Geometry of CLIP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.045027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.045027Z digest=sha256:5f3f90c958f72721597e37f690a2da89695ab68f815850fe0ccf4cfc87663dfb

Observation 4d502446-6440-47de-b3a7-37f7735ac5f9 · outbound

This paper cites Does CLIP Bind Concepts? Probing Compositionality in Large Image Models.

On the rankability of visual embeddings Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.071811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.071811Z digest=sha256:c6d661d0a3fc137f8fe2ed8e02cdfa1305a53376f3a2c7a227c866e6dbe60c40

Observation 2624fd97-d6cf-4533-b4f8-02e1e764d3dd · outbound

This paper cites Align before fuse: Vision and language representation learning with mo- mentum distillation.

On the rankability of visual embeddings Align before fuse: Vision and language representation learning with mo- mentum distillation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.494930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.171974Z digest=sha256:ebdf01bbf52d706b3824963f7f7a1bd7cfac17853effe0f19e2abab171572e26

Observation bf413c30-217d-4123-95e4-6fe4efb58ffe · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation.

On the rankability of visual embeddings BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.478419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.232772Z digest=sha256:52c087811139d9f21091e7048bf5d4fd5268de55fa6da0ac844121b578b8a273

Observation f366b8f4-65fd-4e5b-8ae5-52c09621ca3b · outbound

This paper cites OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression.

On the rankability of visual embeddings OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.956760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.302801Z digest=sha256:086a2e0b9a767da22bebd71cd4176ade8d7898c95fd3de2c4e6d8392a1ae11df

Observation 00efac1b-a70d-4467-b8f4-0ba5c81b0c6d · outbound

This paper cites CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model.

On the rankability of visual embeddings CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.933505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.386439Z digest=sha256:f61533b6f75678526457184ac1e5063a124a26b39ecab78453ea8f736f9e3dd7

Observation ffb1b424-5803-447c-983e-274b6ec42009 · outbound

This paper cites Beyond comparing image pairs: Setwise active learning for relative attributes.

On the rankability of visual embeddings Beyond comparing image pairs: Setwise active learning for relative attributes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.462867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.467629Z digest=sha256:8f70678a3d8301a7c06baade9026dd6a0c8e5a9e77b88d3285610140750603a7

Observation 9cdd71ea-a5ce-4033-93fd-6d02518fd554 · outbound

This paper cites CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification.

On the rankability of visual embeddings CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.500131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.500131Z digest=sha256:c6c4388f7fc40313a927a0ba5981c357871412efdd12b038798d581bc8ccb009

Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.511161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.511161Z digest=sha256:6cc698ba27b20cb1789fbbfcb8b176b58715c64bf150a5600c2dbecac181b171

Observation ead6e43e-4d76-4484-8b92-a25d7b2ded35 · outbound

This paper cites A V A: A Large-Scale Database for Aesthetic Visual Analysis.

On the rankability of visual embeddings A V A: A Large-Scale Database for Aesthetic Visual Analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.516129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.516129Z digest=sha256:b1a390119be8167ed4fc8010c18fbc36fa1700e508e915b23d3dcb3466aefacf

Observation 55c49fff-8449-4320-8028-83eba01abd3f · outbound

This paper cites A metric learning reality check.

On the rankability of visual embeddings A metric learning reality check

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.441046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.520788Z digest=sha256:20f3467fba25d83259d6d1e439a5af287887b1fc88eed68895571665d365a5df

Observation f8ee04c5-2aa8-46e1-be7d-9572b7fa1707 · outbound

This paper cites Parts of Speech-Grounded Subspaces in Vision-Language Models.

On the rankability of visual embeddings Parts of Speech-Grounded Subspaces in Vision-Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.533367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.526319Z digest=sha256:ab5d133f170afa98779def79aeca115ab170d7a88db335443934d4201d808b05

Observation 18e7cfd6-4087-4ef9-beec-b45a1a211e6a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

On the rankability of visual embeddings DINOv2: Learning Robust Visual Features without Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.531605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.531605Z digest=sha256:9050e3e2d716227c92104b252f40f120cc6b39a11023410fba831ded2a13dbfc

Observation bd3a1569-da58-45a3-826f-5ab8b6110eaf · outbound

This paper cites Teaching CLIP to Count to Ten.

On the rankability of visual embeddings Teaching CLIP to Count to Ten

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.537121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.537121Z digest=sha256:57ee61f4f64c7b1f4a93e208f0300046aba34b90268df21449d1d1fdfae7136f

Observation 75fa21f4-3934-4458-8b3d-b3fd6aaebe20 · outbound

This paper cites Dating Historical Color Images.

On the rankability of visual embeddings Dating Historical Color Images

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.542600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.542600Z digest=sha256:bedbb97abd66a6446321ab03d295bf7e55aecdb981192b53aa58a498efd99266

Observation 5e8fa665-8c00-4a0b-b22f-f10365a35b98 · outbound

This paper cites Relative Attributes.

On the rankability of visual embeddings Relative Attributes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.548487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.548487Z digest=sha256:f304d73c2d643cea7bf88304d951883fb4bb4c8aa8344a3f1e1e956868a324d4

Observation 0babb719-b950-468c-a839-5b56f2279032 · outbound

This paper cites HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports.

On the rankability of visual embeddings HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.422955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.553683Z digest=sha256:03c7bb80683988582cfab463de49557f671554197be29b921565caf2696cd13f

Observation a0e88aef-b4aa-4aff-90ec-4a7203a7dd5e · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.559058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.559058Z digest=sha256:f86907bfdec3454246b325fa62d291552da98149deaec9088ff19203cebc8655

Observation 025ec07b-d7e3-459d-99a9-9abf9a1a44e9 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Super- vision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.404122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.564088Z digest=sha256:3ab1975533ca41bd5732ba3a451792afaa5ce1bfdc84f7c19ac1a13a1d5a021e

Observation e2d4c7c3-cb2a-4eed-bf32-218603569642 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

On the rankability of visual embeddings Steering Llama 2 via Contrastive Activation Addition

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.569573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.569573Z digest=sha256:1e332edbef7e6546402f7fdf1b0ead081ab0c1d78c748b22f95bf5096315cd1f

Observation fe3b8e31-c769-4e73-ae32-dd0e949665a4 · outbound

This paper cites Finetuning CLIP to Reason about Pairwise Differences.

On the rankability of visual embeddings Finetuning CLIP to Reason about Pairwise Differences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.575686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.575686Z digest=sha256:6ce170c7e41067ec2f748175994948e2799dc449e3ff9e2e602cdb55346993b7

Observation 3fd03dbe-f488-4d9f-be8d-8c43e02a1b47 · outbound

This paper cites Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification.

On the rankability of visual embeddings Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification

Reference 49

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T20:11:52.391746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.583085Z digest=sha256:592f74d9bd1860371d499e5809776a3b3d43c54c6b78e3057d5760355cce8cf5

Observation bae4703a-499c-46ee-8035-7a7ab5c85cd7 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.386329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.588973Z digest=sha256:87afa5c90c88942d648b8e6eb73fc98d37a92469550f073d810cf4d9fbbe0d79

Observation 3238f178-81da-4011-99c2-d6fe4be79ac2 · outbound

This paper cites Linear Spaces of Meanings: Compositional Structures in Vision- Language Models.

On the rankability of visual embeddings Linear Spaces of Meanings: Compositional Structures in Vision- Language Models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.368075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.601038Z digest=sha256:49b3aa5f0ca6cd3bdb167816eea5b8b743516cdc035ae532edd2b0b23eddca98

Observation 1b0cdd12-01f9-4d10-9361-19f08ebb1ee4 · outbound

This paper cites Deep image prior.

On the rankability of visual embeddings Deep image prior

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.348792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.607467Z digest=sha256:88eaeccd3f70bd6e692d4422e94f9bcf35934e6080dba14ca7933ede8c7de775

Observation e4fe7ec6-2185-47a2-af26-0d7a5a3c787e · outbound

This paper cites BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION.

On the rankability of visual embeddings BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.330263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.612865Z digest=sha256:5bc239c69822bcbcb93ffad29bafbfcb8204c51fffe4c46df727b0c25b2f40f8

Observation 4230b46f-7018-4a08-8e07-12086fe5af61 · outbound

This paper cites Intermediate Layer Classifiers for OOD Generalization.

On the rankability of visual embeddings Intermediate Layer Classifiers for OOD Generalization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.308628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.617981Z digest=sha256:f8359aca633e8ab2c30fdfdb4d2683c3e389d8c9d97a17ded527c794dc8f1e8b

Observation 9247635d-3ab4-46b7-847f-bf53ff49548e · outbound

This paper cites Order-Embeddings of Images and Language.

On the rankability of visual embeddings Order-Embeddings of Images and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.623271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.623271Z digest=sha256:b3b595e1aafb0bd04bfae0dfb3fcb927e352ccff795efb93604c1d141f18f6c5

Observation 6b5e8000-421a-459c-a408-8b76f589aed6 · outbound

This paper cites Exploring CLIP for Assessing the Look and Feel of Images.

On the rankability of visual embeddings Exploring CLIP for Assessing the Look and Feel of Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.628732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.628732Z digest=sha256:2732aef48fb17354fa1c883ef3c811395c56faf2f005204994dd9ccb43155ea3

Observation 936ac420-362e-4fbf-8e94-e91a8469f375 · outbound

This paper cites Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification.

On the rankability of visual embeddings Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.290182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.634171Z digest=sha256:c9b69fb70eba141c6f18c598126a21b61b728496561667d7602fa8962392e972

Observation e4732314-04f8-4044-ba8f-e92acfb33c58 · outbound

This paper cites Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification.

On the rankability of visual embeddings Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.271360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.640208Z digest=sha256:be1ea3fac8f6f30262028ede81b641769bb848113a9e5b510542184f21ba732c

Observation be34882b-7800-4c43-bc56-e814d8e5511c · outbound

This paper cites Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere.

On the rankability of visual embeddings Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.646034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.646034Z digest=sha256:db3d1cb2a450f56c2c32ee21f47c4b803292dc1b52302a157c5bc862b3762ead

Observation 6f1bb15f-0122-4804-9671-9633a12c7e06 · outbound

This paper cites Disentangled representation learning.

On the rankability of visual embeddings Disentangled representation learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.252755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.651508Z digest=sha256:6f1002a7435da621db46e56154f7b1b1b0bf58166a95fac8b1805e463e0e108d

Observation f0ff0cf9-3c50-4858-9037-a99cb6c6d5d9 · outbound

This paper cites ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders.

On the rankability of visual embeddings ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.656002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.656002Z digest=sha256:5139bd086e4309b6f0d0613977d04d1aee3c9b46c027d864b827b1f3cc8045f1

Observation c36c12cd-79c4-42e7-a276-151ea491f4ee · outbound

This paper cites CLIP Brings Better Features to Visual Aesthetics Learners.

On the rankability of visual embeddings CLIP Brings Better Features to Visual Aesthetics Learners

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.250347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.661258Z digest=sha256:2e1ac6acda7fa02282c3051837bc35192c5448150f896e4ad389f2ea0a78b2fa

Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.667082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.667082Z digest=sha256:dd727771ea363606dfe0b1d5feccbad327481cc8a1ab468a096dbec68d626589

Observation 01c94fe3-87ce-4cef-8ba2-51249241ef8a · outbound

This paper cites Just noticeable differences in visual attributes.

On the rankability of visual embeddings Just noticeable differences in visual attributes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.233523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.673658Z digest=sha256:647de55e3b3c1a7168188dfd83c9033a025d328201d7d0fd89f1349b7e5e0709

Observation 82ad9c1e-42a9-416c-b056-06aac0007fb0 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

On the rankability of visual embeddings CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.216137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.679214Z digest=sha256:3281ee7b4d65f14a9f17664a200eeb9166119f76b19235bbaef517f3a539ec92

Observation db9bfc04-6741-4d31-9d31-7c52aab8b687 · outbound

This paper cites RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP.

On the rankability of visual embeddings RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.200041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.684176Z digest=sha256:4dcc6261a00a1a8670a699c0356f93d568e230d7ea77a185a95ff2b29e73b8b1

Observation 3f14a96a-9b7b-416c-97c0-65bfe82f4afc · outbound

This paper cites Ranking-aware adapter for text-driven image ordering with CLIP.

On the rankability of visual embeddings Ranking-aware adapter for text-driven image ordering with CLIP

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.801956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.689502Z digest=sha256:3ba7d8fb11b8dbdb323755634d15e456dcde78bb5bbca6270cfb3dbb3fccb45a

Observation 3ed8fddf-ba6d-48cb-8c12-f3952ea698e1 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

On the rankability of visual embeddings When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.695389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.695389Z digest=sha256:f91cb100c077b738e5a049aed5e711adc36ead1065d6a9d31e1a8deeae1b3bd6

Observation 68d75df2-e218-4e43-b214-cd4e2f6db13d · outbound

This paper cites Single-Image Crowd Counting via Multi-Column Convolutional Neural Network.

On the rankability of visual embeddings Single-Image Crowd Counting via Multi-Column Convolutional Neural Network

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.701459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.701459Z digest=sha256:aa510fecb4a128be25fede86988e8d1532be754643cf6c2920bfcb2878cdf0c6

Observation ef6fcadf-a217-4378-ae55-e1e091615786 · outbound

This paper cites an unresolved cited work.

On the rankability of visual embeddings Unresolved cited work

Reference 70

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:52.187722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.706634Z digest=sha256:8c822d8896cf48475ecb2338608566c3c0202abbd5b80e92bad414d570186e8c

Observation e330120c-f2ee-478c-809d-c67b5537b483 · outbound

This paper cites Learning Ordinal Relationships for Mid-Level Vision.

On the rankability of visual embeddings Learning Ordinal Relationships for Mid-Level Vision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.183574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.717772Z digest=sha256:321da865a235c41bf187aa8fe8a3a465cd46c5e6701ceea903fd05f22aefe7d5

Observation 73c93712-e3f5-4205-97ab-3869d8d60f4d · outbound

This paper cites Age Progression/Regression by Conditional Adversarial Autoencoder.

On the rankability of visual embeddings Age Progression/Regression by Conditional Adversarial Autoencoder

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.767240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:11:51.712243Z digest=sha256:f91679cafb097521a03b4a7abf7eecc92c6fc5cc0dc18246b41c7f4bce544b2b

Observation ec1d0391-a4a0-4c85-84d5-c2f4ae582f49 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.594734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.594734Z digest=sha256:bd9430310e451a5846f4c51d77d1b453ae33848ce2121965c48d524e4bec940e

Observation 7d03c2f3-6c39-41dc-89c0-7120d1ade950 · outbound

This paper cites DOI: 10.1109/TIP.2020.2967829.

On the rankability of visual embeddings DOI: 10.1109/TIP.2020.2967829

Reference 4056

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.176194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.176194Z digest=sha256:606930ea1a4a8ccb6e13ff6fe8c2a86f887555b3006d7a97421d4b5def7525a0

Pith citing papers

No inbound Pith citation observations are available.