Pith. sign in

Paper Citation Record · LEDGER

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

As of 5 August 2026, this Paper Citation Record lists 100 of 126 outbound references and 1 inbound Pith citation observation for arXiv:2607.18237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18237 v1

Coverage vector

measured 100 of 126 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:43:04.068454Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:43:26.613972Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T11:43:27.703168Z

Reference resolution

100 of 126 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8dc5014c-b66e-4ec7-84ef-5f1398f139f8 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE Transactions on Image Processing, 13(4):600–612, 2004.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Image quality assessment: from error visibility to structural similarity.IEEE Transactions on Image Processing, 13(4):600–612, 2004

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.320824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.320824Z digest=sha256:88c1fefe48336ef0dcd87046c7f1b47a850073eb9ce86ebf498236c08d593953

Observation ad0ac39b-0119-4ba1-9639-d0581e83955c · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric The unreasonable effectiveness of deep features as a perceptual metric

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.395626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.395626Z digest=sha256:bdf7bdcd9497782dfc2bb07b48c5814f52748c6e9eafa11e6128129b531ad6ca

Observation d72e7b48-1f23-48ed-a4e0-9ef9d235f3cc · outbound

This paper cites DreamSim: Learning new dimensions of human visual similarity using synthetic data.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric DreamSim: Learning new dimensions of human visual similarity using synthetic data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.489385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.489385Z digest=sha256:6bc3244f3531fde40706fc5a7813a4e02a633d2bdeed96355bfa6847f2ebc20e

Observation 081c99c3-7c97-430b-b06d-db048ae83c39 · outbound

This paper cites Features of similarity.Psychological Review, 84(4):327–352, 1977.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Features of similarity.Psychological Review, 84(4):327–352, 1977

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.580404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.580404Z digest=sha256:c88215340d355fc68110cf6d169c0d375d2b3dd3c05aee18f0592cf22c4174c8

Observation 5abb6494-328a-4871-9049-26b3aa513d8b · outbound

This paper cites Respects for similarity.Psychological Review, 100(2):254–278, 1993.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Respects for similarity.Psychological Review, 100(2):254–278, 1993

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.674454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.674454Z digest=sha256:30e8467ff6bd86a36e12d626f825ce91b1174b11c3713f9533e4edd4ab201850

Observation 18564c58-40bc-4e3e-96b1-14576272c6aa · outbound

This paper cites FSIM: A feature similarity index for image quality assessment.IEEE Transactions on Image Processing, 20(8):2378–2386, 2011.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric FSIM: A feature similarity index for image quality assessment.IEEE Transactions on Image Processing, 20(8):2378–2386, 2011

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.761841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.761841Z digest=sha256:346530be08d42924072ebd4e7470373f5f28f41a880e3d9657921b10db44ecfd

Observation 7aafc947-5ca6-44a0-b060-eca9fa8bec2e · outbound

This paper cites HDR-VDP-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions.ACM Transactions on Graphics (Proc.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric HDR-VDP-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions.ACM Transactions on Graphics (Proc

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.826536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.826536Z digest=sha256:a24cbf0941a25f1efa2fbdcb5128b89a0ffc2534e4508fab548c1c4bc77d7add

Observation 0fdbb787-7bd3-4a29-a8d0-789fac99db71 · outbound

This paper cites GeneCIS: A benchmark for general conditional image similarity.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric GeneCIS: A benchmark for general conditional image similarity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.896997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.896997Z digest=sha256:61c4fecffa172f6090c15aa625e3e03db9069351ad9c13ecdc2ca8148badc742

Observation e866951c-e765-4f68-be28-5bb3b0ca250f · outbound

This paper cites FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:53.983507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:53.983507Z digest=sha256:f67c6ad04cdd882db2145fd3022ca7aeebf9a793dad64821243c277bde15d8e2

Observation 084f9d75-3c5e-4095-bc25-415c17a4a686 · outbound

This paper cites Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.057921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.057921Z digest=sha256:339729f86eae1da03a47c5509ff98dea579ef0ca3ddab5c1a7061b8065b9df13

Observation b4cdf446-56b8-47eb-b3a9-c415239d2327 · outbound

This paper cites THINGS: A database of 1,854 object concepts and more than 26,000 naturalistic object images.PLOS ONE, 14(10):e0223792, 2019.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric THINGS: A database of 1,854 object concepts and more than 26,000 naturalistic object images.PLOS ONE, 14(10):e0223792, 2019

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.136636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.136636Z digest=sha256:a85e8c521f2330da2ee8f9265a541becb674873bf365347842b3e9ffb9310f8f

Observation 4f2682db-d7db-47f9-8042-768ef358c7dc · outbound

This paper cites GPT-4V(ision) system card.OpenAI Technical Report, 2023.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric GPT-4V(ision) system card.OpenAI Technical Report, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.241360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.241360Z digest=sha256:01097b70a5f3a1197d1ac0a929fa006d6ddc8203de3064fe45732403b678c609

Observation 6891513a-4177-4ddc-86a7-ae453f3c3854 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.321607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.321607Z digest=sha256:2ae6f1369bb135b117edc0f222d7054f5da466f49303f9066877d592a0565c5a

Observation 5c98f41c-fa32-4afa-84c8-f293784391f2 · outbound

This paper cites Qwen3-VL Technical Report.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Qwen3-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.413463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.413463Z digest=sha256:1609579c85d370b2984db2c5e6e6d5d4d29bb0be8e47d5ca201db2ea9bacb926

Observation 5cf9da9e-3365-4a08-89bb-7f733632d7b9 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.466808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.466808Z digest=sha256:22fa63f8561700a8b2117fb3e5fba22b12e2a6c5b73c3ddc25bd13736f1162ad

Observation b05a11f6-5641-463a-bd23-f00f2b95190d · outbound

This paper cites Qwen3-VL-Embedding and Qwen3-VL- Reranker: A unified framework for state-of-the-art multimodal retrieval and ranking.arXiv preprint, 2026.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Qwen3-VL-Embedding and Qwen3-VL- Reranker: A unified framework for state-of-the-art multimodal retrieval and ranking.arXiv preprint, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.607195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.607195Z digest=sha256:594df5164ea96a778cf63335a0d0d5909cf4d53f203bf0faf92132302e8e4ea7

Observation c8655d58-0f1c-4d42-a28e-85b30032bd61 · outbound

This paper cites ImageNet classification with deep convolutional neural networks.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric ImageNet classification with deep convolutional neural networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.664255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.664255Z digest=sha256:56478e85c0fd0f027736bae466affc509fc6c15eb22ded1835f2cfcb8596c16b

Observation 0a9ea458-a14f-454a-a445-a4755c8f0bc8 · outbound

This paper cites Very deep convolutional networks for large-scale image recognition.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Very deep convolutional networks for large-scale image recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.754588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.754588Z digest=sha256:134d2cc277c8278ace591f4f7847fba359551404be9a59f6882d58bd3a8e8f8a

Observation 488b3466-ecb6-4dad-b85a-b7cca1dd9ccd · outbound

This paper cites Deep residual learning for image recognition.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Deep residual learning for image recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.836815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.836815Z digest=sha256:3d8aaf68f14316937441a7270cc3129dbfa4e020d6b838b7c01bd0e42678a676

Observation 72f09c5f-f7b1-413c-9818-cff1cb957a83 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric An image is worth 16x16 words: Transformers for image recognition at scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:54.955281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:54.955281Z digest=sha256:1f45d1e370c4b30c6ee93ec9dad044c6a8842a6c7d73516f3e6cadf7be6779ab

Observation 3f1a9ec6-4828-4f7b-8152-fd132c2b5c08 · outbound

This paper cites Learning transferable visual models from natural language supervision.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Learning transferable visual models from natural language supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.033336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.033336Z digest=sha256:489b8a5c6114f8092f378a4a53f7ac9cb115f2adf16487bf6c9fff04a0d5e2d2

Observation f8129422-0d3c-4289-a6c5-e758dcd89b3e · outbound

This paper cites DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research (TMLR), 2024.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric DINOv2: Learning robust visual features without supervision.Transactions on Machine Learning Research (TMLR), 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.112667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.112667Z digest=sha256:157ed87f9fa31ec5d0284ff74775da305c576e9bdc7a78830e9b3f9536ce2ce4

Observation 0454be89-7089-4b22-8986-e390403c852b · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Perceptual losses for real-time style transfer and super-resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.194507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.194507Z digest=sha256:a04d8bebba3179506303f27525d4ed61c00af48ff336a5c0cb8bc53c070fb1fe

Observation c12744ba-0c50-4dc2-a2c2-c9a6ea4c9f46 · outbound

This paper cites Image style transfer using convolutional neural networks.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Image style transfer using convolutional neural networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.280869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.280869Z digest=sha256:96ca97ef36841c7d760742d0040ce135b8c1b5943aee9a68e1d9605cf861c699

Observation 927175c8-bfa4-46d8-bc0d-e899ddfa27b4 · outbound

This paper cites Understanding and simplifying perceptual distances.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Understanding and simplifying perceptual distances

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.371081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.371081Z digest=sha256:1c2b1e94cfb1ebe05c879eea736b99c1492fbc5dada107f4932ff4de6c6b84e5

Observation b7080aab-72e5-4d32-9bf5-77b40d7207b1 · outbound

This paper cites Vlic: Vision-language models as perceptual judges for human-aligned image compression.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Vlic: Vision-language models as perceptual judges for human-aligned image compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.432729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.432729Z digest=sha256:9c17732f0a42e0f3bf6e5c71a871b392fbd89eb07e05f0435eaeed475fa58036

Observation 50476618-fc88-4635-a2fd-8c2a3785a941 · outbound

This paper cites PieAPP: Perceptual image-error assess- ment through pairwise preference.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric PieAPP: Perceptual image-error assess- ment through pairwise preference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.530735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.530735Z digest=sha256:3fc818d58cca9edcdd80935578cf31f4872823e76132289a87a6ecc259110fc9

Observation 00ea3663-c251-4015-92cd-526226f9f6c8 · outbound

This paper cites Image quality assessment: Unifying structure and texture similarity.IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 44(5):2567–2581, 2020.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Image quality assessment: Unifying structure and texture similarity.IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 44(5):2567–2581, 2020

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.618632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.618632Z digest=sha256:2e1fa3a8c52951faf982762268811015c128ec4eb9f84a18a0ebf70c0d87d393

Observation 2a4614ee-9868-497f-aa44-2b752b15fe5d · outbound

This paper cites Human alignment of neural network representations.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Human alignment of neural network representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.687313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.687313Z digest=sha256:ee06d5dfe13f51338788519c06d910c918fa00f588b0bbe57af98252e3812ee4

Observation 9019f240-3579-4fe4-b9a5-c7f6068175bc · outbound

This paper cites Shift-tolerant perceptual similarity metric.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Shift-tolerant perceptual similarity metric

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.765243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.765243Z digest=sha256:3bd69216c96f4e27eaec1db7ce69efcb3e6631329f213905f42b69d1e65a8191

Observation 39420102-73cd-4fc0-9f0c-dc59a7e41211 · outbound

This paper cites LipSim: A provably robust perceptual similarity metric.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric LipSim: A provably robust perceptual similarity metric

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.829960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.829960Z digest=sha256:52d19c8e3c6105fe9a1769de671f9a2c69ff102049404e07073dee960f44dec8

Observation f2b8669c-d59c-4bf6-82fd-2460308c231e · outbound

This paper cites E-LPIPS: Robust Perceptual Image Similarity via Random Transformation Ensembles.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric E-LPIPS: Robust Perceptual Image Similarity via Random Transformation Ensembles

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.882093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.882093Z digest=sha256:31a2de82e5841fc7e8b57bfee6f00a4eaef9e25b48d385eea4209508f8dc11f1

Observation 3686a8d0-369f-4379-bcbb-85588218ae04 · outbound

This paper cites A similarity measure for illustration style.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric A similarity measure for illustration style

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:55.995922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:55.995922Z digest=sha256:788be8c16f8f123c6827d86b82077c46d625e09bfae30d99ef6be4d6ed6cf045

Observation d1375da2-170e-4127-aecc-7b8d15144729 · outbound

This paper cites Measuring style similarity in diffusion models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Measuring style similarity in diffusion models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.057269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.057269Z digest=sha256:f926ed941d6d42fb35933f7dddbbb944a97f056c5cd4ec48400added4e8c37b4

Observation c659abb5-99d6-407f-8caa-c1c2a9de8087 · outbound

This paper cites Relational visual similarity.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Relational visual similarity

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.120843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.120843Z digest=sha256:71853acf7da92980fae1c226a692d8425bbcfa40fd74e7fa9619d40f361e22da

Observation f08eb537-a75e-4a2e-9991-5f319e2cbc72 · outbound

This paper cites ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.203684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.203684Z digest=sha256:077552b730bcabc1f92a5cdc57a9a50e7e4daf20102e0cb1b9eec6462a2250fc

Observation 6e10fb11-c540-4bfe-918c-a6513c5907d8 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.326331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.326331Z digest=sha256:bc45ec6d3c4041e27ed5d3782539e9d3c4b217ba7a6f59ed8dfb4b5e408c1e34

Observation 70d65d67-80fe-487c-90d4-79b5dd8b1399 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.376202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.376202Z digest=sha256:6de778ce3c0d95304c5f424cc8afb09307bdcbdef169992c9029e6142c86b2ad

Observation 02c44e49-5639-40eb-99e0-e3299155a138 · outbound

This paper cites Editreward: A human- aligned reward model for instruction-guided image editing.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Editreward: A human- aligned reward model for instruction-guided image editing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.466663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.466663Z digest=sha256:4507e9677b270d4d3bf042f9bba8b99defa610fdcefb901c0c7d4b5a45403e93

Observation 966961fc-d656-4e71-abe2-7cbe178bb2d8 · outbound

This paper cites Multiview triplet embedding: Learning attributes in multiple maps.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Multiview triplet embedding: Learning attributes in multiple maps

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.540947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.540947Z digest=sha256:3bcd02ce5f3a296804f760c73840626af27396396246d0e83b455958299dd88d

Observation 28652472-8002-4b0c-926b-860ef16ea6a1 · outbound

This paper cites Conditional similarity networks.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Conditional similarity networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.591116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.591116Z digest=sha256:ea55f89f36b6b683d5a1279b536b78498a3841aa9e86022b1c26e961639da69c

Observation 4ec6f53f-8218-486d-814c-3799c4dcffb2 · outbound

This paper cites Learning type-aware embeddings for fashion compatibility.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Learning type-aware embeddings for fashion compatibility

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.677880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.677880Z digest=sha256:894a4aaf5e389c8c71ebcc0d1ec4e13c59578fe1c2c90dd7a1b687c68da57b4d

Observation da2aa78c-e5d1-4181-84be-12321b230a37 · outbound

This paper cites Cooperative Embeddings for Instance, Attribute and Category Retrieval.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Cooperative Embeddings for Instance, Attribute and Category Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.761219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.761219Z digest=sha256:020841105ab3410a2c485b53778408ce519c57dc8b859cd24c8346524433e2ed

Observation 5955305d-87a7-4a6d-bf8c-11cb0c48222e · outbound

This paper cites Learning similarity conditions without explicit supervision.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Learning similarity conditions without explicit supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.850230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.850230Z digest=sha256:739ee8b1c6e4ba7df4c557b8a3d07af055cf4d54b396d98d3e7f7bd2c24af206

Observation f5b7c9ac-320b-4247-b533-eaa5bba7d40f · outbound

This paper cites Training-free conditional image embedding framework leveraging large vision language models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Training-free conditional image embedding framework leveraging large vision language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.894850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.894850Z digest=sha256:cee94038ddbe3b07dc8105e18feb86e3f54b2d16cd8dd63fe9a165488f603ec6

Observation 175e6537-4256-43e6-84ce-89cb91fa9897 · outbound

This paper cites Highlighting what matters: Promptable embeddings for attribute-focused image retrieval.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Highlighting what matters: Promptable embeddings for attribute-focused image retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:56.970914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:56.970914Z digest=sha256:0d0ff5e096f06ca10bdd1d61b75b5ae65ac87d18ca714e36e0d9ff7488882984

Observation 67a2f29a-81f2-4212-b6d8-f6b768e4d95d · outbound

This paper cites Towards text-guided attribute-disentangled multimodal representation learning.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Towards text-guided attribute-disentangled multimodal representation learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.031424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.031424Z digest=sha256:857c615f3c534c51ae61c03dd71a235eda80a05ba118fd07ca5168c83352c7ea

Observation 93473ce5-bf2c-437c-b3e5-42d6ff96e2dd · outbound

This paper cites Open ad-hoc categorization with contextualized feature learning.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Open ad-hoc categorization with contextualized feature learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.093953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.093953Z digest=sha256:45e25809f509f12202ccc76189f40c09570a223478e327b0511d37a4fc0ef5d9

Observation d6926112-2b70-4479-92d8-7e048ad99725 · outbound

This paper cites Image clustering conditioned on text criteria.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Image clustering conditioned on text criteria

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.156811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.156811Z digest=sha256:5ebbff87e8814605ff5cdc2bdfcb88cb15ae1822cc57c3ccde7787d464ecdf25

Observation 130a5a0b-d2bd-4229-865d-2a881a330d78 · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Composing text and image for image retrieval-an empirical odyssey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.224751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.224751Z digest=sha256:c0428a04d188fbace1044e672c9f01ccbf95a750f61909d6b5108f5a92f98d89

Observation b7cc5668-677c-4f67-afa0-ab77d3d62934 · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Image retrieval on real-life images with pre-trained vision-and-language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.332809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.332809Z digest=sha256:2eb24d84d84e16cabf3b272107f884c4e1a3550e30467429c7ca98005deac01e

Observation f45de66c-8d84-4d19-9ae8-2fa0f6b1d1d4 · outbound

This paper cites Pic2Word: Mapping pictures to words for zero-shot composed image retrieval.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Pic2Word: Mapping pictures to words for zero-shot composed image retrieval

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.412517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.412517Z digest=sha256:7af7e5a19dbb07503a2baf848c9c8b6b5cc3d0466adf90d715c9cc0d9443d2ee

Observation 8dd60f0c-aef9-4e79-ac9c-c4639a2a3369 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Zero-shot composed image retrieval with textual inversion

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.461199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.461199Z digest=sha256:70b7dd29edd6075297392d897f23e1c95ebdfd3f9e19c299c76a8941a2ec48f2

Observation 645ea2f7-b214-46c8-bef3-fa3e07a14d00 · outbound

This paper cites Sigmoid loss for language image pre-training.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Sigmoid loss for language image pre-training

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.539094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.539094Z digest=sha256:aebeacb49c61eddabe9da20c7c7019c5e4fcdedf4cd2d53f3282395ef4e0a4e9

Observation 5c3d6f6c-7e32-46ca-bc03-c6c96396347b · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Reproducible scaling laws for contrastive language-image learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.631959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.631959Z digest=sha256:49b564e44dd87e96732cfe325e1ac48f892df559d119f6a06173a55fecd73218

Observation d8c60eae-40d8-4ef1-ad02-d7f9515680ed · outbound

This paper cites BLIP: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric BLIP: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.732384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.732384Z digest=sha256:ebd6f369ee28525e9c5147c4adc0a87ee126acecbba3604c0f6c5f23e5c9ce29

Observation 89bd14d8-46e9-4e7f-af89-cf5fc56195e7 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric BLIP-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.897656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.897656Z digest=sha256:d8032bfa8bb9f1d6bc4863940e92e901e6eb9d0489b111a115917138938aad83

Observation 4c686426-99a3-43a8-a642-b30810092525 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.057377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.057377Z digest=sha256:cb2ea73ca4e23301b1e29f059676fac7795b7ff0f20f30803541aaa1ae921b54

Observation 71bf44e3-ebfe-4de4-bcd4-c4898bacf59d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.186290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.186290Z digest=sha256:e2843e4eb1c13e82a8a08d3a2cdd2a3c803446ac1aeda46b97f5997c489a881c

Observation 13734fc6-08d5-4800-a500-2f100b266e64 · outbound

This paper cites Qwen3.5: Unified vision-language foundation with early-fusion multimodal training.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Qwen3.5: Unified vision-language foundation with early-fusion multimodal training

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.285881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.285881Z digest=sha256:2005603622750182267d548e285ea2ee000f3fe23267b4fd91ccca3ca0f381da

Observation 3c70b9e1-c968-4ffa-81db-0bcc6a859743 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Flamingo: a visual language model for few-shot learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.438877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.438877Z digest=sha256:4484fa2962ab5cd86f133069527e798a28d64440ea7c6a845f3b91a277674a40

Observation 9f34bec9-4b25-4d99-b062-34e97db517f6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric LLaVA-OneVision: Easy Visual Task Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.563681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.563681Z digest=sha256:4d13789b15e89d791b1a35fd92a5ca5f8c992a471839ebd47a1a7e33eeb8d6f5

Observation 743802ab-1fcd-41c9-9484-12253816ced1 · outbound

This paper cites Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.667495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.667495Z digest=sha256:5b44377aab4da5de446fed76bd8141b4376743d925272c4c02165b87f05eab1c

Observation 3180a520-1559-4d03-a37b-b98783286db5 · outbound

This paper cites VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.671408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.671408Z digest=sha256:8953ba890c19292585fefce090d6e4c5fa2518b62d4e31fc208ca8d1932aa10a

Observation 0b8bbb34-8a38-4ddd-9730-8cd48d5abea2 · outbound

This paper cites FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.828756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.828756Z digest=sha256:be5bc18c383edc276fb53297dd339aa17ce6a1f11eb4612c7e0052f670373f14

Observation a4694518-7b44-4ec7-b880-719ef5e9723f · outbound

This paper cites VLM2Vec: Training vision-language models for massive multimodal embedding tasks.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric VLM2Vec: Training vision-language models for massive multimodal embedding tasks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:58.938966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:58.938966Z digest=sha256:8b17c62da47d962de4fb8c3ad5ac9353d065d74c9241c164e8317c4c36fa3bd3

Observation ea80c254-ea72-43b4-95a6-c857450105ae · outbound

This paper cites NV-Embed: Improved techniques for training LLMs as generalist embedding models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric NV-Embed: Improved techniques for training LLMs as generalist embedding models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.105369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.105369Z digest=sha256:bd91c99014a61f19d6f580375db8703217c50316745a7cd1e7fa6c6fe3a7bfd7

Observation db527fbe-bb7d-4ae7-af44-b75c6666e6a8 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.207367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.207367Z digest=sha256:54df6ffb5398e0f1ab00257b739a8d61040a44d07a55ee9a89881dda36136aee

Observation 947b6cfd-089b-43c7-b31f-6e8a2d5d0a7e · outbound

This paper cites Steerable Visual Representations.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Steerable Visual Representations

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.298378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.298378Z digest=sha256:a38f8e17767057f7f6073f25e8ed48f74b4557e6b21b443cac949a20a50c8b77

Observation 19dd8d8b-3b74-4101-943d-a3da6aba1b99 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.364828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.364828Z digest=sha256:768135deb8459144cd03bd05c3bebc0069cd089993ce512cc4894bfe30ef784a

Observation 3567e077-226a-4a3a-8331-cd30ac8f02a9 · outbound

This paper cites The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.International Journal of Computer Vision (IJCV), 2020.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.International Journal of Computer Vision (IJCV), 2020

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.562745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.562745Z digest=sha256:5919fff8e2c841b172dd2fe085c569a80310a03c1e3f09a825c24989ed4d9a51

Observation cd6e6005-888e-4c53-b337-1664986ccb7c · outbound

This paper cites Improved artgan for conditional synthesis of natural image and artwork.IEEE Transactions on Image Processing, 28(1):394–409, 2019.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Improved artgan for conditional synthesis of natural image and artwork.IEEE Transactions on Image Processing, 28(1):394–409, 2019

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.674854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.674854Z digest=sha256:1ab3b884d523525a6398700d10d5feafd07da2e3abd2fd4651b1afcb760109f1

Observation 75bca747-fe7a-400a-a5bc-689eae9b51a4 · outbound

This paper cites Learning an image editing model without image editing pairs.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Learning an image editing model without image editing pairs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.819099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.819099Z digest=sha256:d202f10d7cfab9cc5e9508916add2fc622a254cbf8c6998a94d6c4c98db93b9d

Observation 9e6d918a-5813-46c7-a62a-746d850e083e · outbound

This paper cites Dual-Process Image Generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Dual-Process Image Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:59.951743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:59.951743Z digest=sha256:3954ead24fdca92e1aa5db1b26abc505eb8a5741e8b9ca88f18d3b1a5e4275e9

Observation 5f9655e1-a833-4d0a-be80-56c9e3e76e42 · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric DanceGRPO: Unleashing GRPO on Visual Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:00.107495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:00.107495Z digest=sha256:b3e90c1a904f3cec048fb13dcc47bc38839af36ef595a19aae91133973b599f5

Observation 38b2b3c5-4fb6-4b48-a020-01d129659a8c · outbound

This paper cites A multimodal automated interpretability agent.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric A multimodal automated interpretability agent

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:00.237857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:00.237857Z digest=sha256:7a530ff0963aa43c65d4163f8bfb06d731611f15bb5a9f1bfdf4ca954ea1b8c2

Observation 9ad88404-ad64-4330-9c1f-d457c0dac327 · outbound

This paper cites One-step is enough: Sparse autoencoders for text-to-image diffusion models.arXiv preprint arXiv:2410.22366, 2024.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric One-step is enough: Sparse autoencoders for text-to-image diffusion models.arXiv preprint arXiv:2410.22366, 2024

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:00.387859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:00.387859Z digest=sha256:c537c48d3a2f246ce410093e3386fff39c0d62d51c65c35ae490368091427331

Observation 9a23a057-9712-4d2b-a529-25e2000c94f6 · outbound

This paper cites Qwen3 Technical Report.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Qwen3 Technical Report

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:00.512691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:00.512691Z digest=sha256:2e11df3c5b8d9d2b86db48744e7088f5641941b20a8f02cacecd9eb5deaf8492

Observation fc5fad79-19ee-4883-813b-335d1925ca3b · outbound

This paper cites FLUX.1: Open-weight rectified flow transformers for text-to-image generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric FLUX.1: Open-weight rectified flow transformers for text-to-image generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:00.643897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:00.643897Z digest=sha256:002010bcaf6213a23cd6fa65c70518d3e6284b5f68e2d0a2d74020d52e560ce8

Observation 615b92af-591a-4b71-ac93-08b6b01c1bfe · outbound

This paper cites Geditbench v2: A human-aligned benchmark for general image editing.arXiv preprint arXiv:2603.28547, 2026.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Geditbench v2: A human-aligned benchmark for general image editing.arXiv preprint arXiv:2603.28547, 2026

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:00.768585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:00.768585Z digest=sha256:58ba64ba6fd9b0eeda791495bffbe899cc085987b937b369a21e966415d6b93c

Observation c99bb2f7-a827-43fe-9832-1738a2878612 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Emerging Properties in Unified Multimodal Pretraining

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:00.920862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:00.920862Z digest=sha256:ec5196830d54652f4403dd067d038f4cce499d376d4cc16fec3435bb96ff2386

Observation 3977a8ec-ae3a-49bc-b9b0-9e896efa33aa · outbound

This paper cites FLUX.2: Frontier visual intelligence.https://bfl.ai/blog/flux-2, 2025.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric FLUX.2: Frontier visual intelligence.https://bfl.ai/blog/flux-2, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:01.050501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:01.050501Z digest=sha256:11d2af2971e5fb815ed2715027a90cf580525fbb13b95b49a656f9de1a7a79c6

Observation 27429766-5990-48f7-bd3e-34c7de2fc427 · outbound

This paper cites LongCat-Image Technical Report.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric LongCat-Image Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:01.150238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:01.150238Z digest=sha256:af50bc7122a25c93f39ff78f5ff0fe336171db9e62b2c801a2c8c80577ddae4c

Observation 66bee9c5-1416-434d-a659-7201fa5b692c · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:01.255956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:01.255956Z digest=sha256:0ae2de1ed033ce02e9eddb0f88a3f9e257feb8087c42efef612218dadc8da7de

Observation b2530244-0200-437d-9e6b-fbcb7ee407a4 · outbound

This paper cites Qwen-Image-Edit: Image editing with higher quality and efficiency.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Qwen-Image-Edit: Image editing with higher quality and efficiency

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:01.464881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:01.464881Z digest=sha256:f527199689b7ab5ffff6c7524dfec0b1d9c857bc9de92ab569c43b2c06156133

Observation bd1f7265-f2af-492f-a7be-dbf4c13d24f8 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Step1X-Edit: A Practical Framework for General Image Editing

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:01.644727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:01.644727Z digest=sha256:8435a8bed222672c49f36f6cf5ab1dfb220a266c49959b7ff595e98dafcaeaed

Observation 91551bca-1c00-40ac-b442-2b8fc73ae516 · outbound

This paper cites Scaling in-the-wild training for diffusion-based illumination harmonization and editing by imposing consistent light transport.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Scaling in-the-wild training for diffusion-based illumination harmonization and editing by imposing consistent light transport

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:01.832622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:01.832622Z digest=sha256:d0ee8b554886ed15b66bbd33c849a92cb0e2459d39f4c970fbee2850c7cfdf7b

Observation 70f8c7a1-191d-4bd2-8379-a26d68c1e3fb · outbound

This paper cites DreamLight: Towards Harmonious and Consistent Image Relighting.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric DreamLight: Towards Harmonious and Consistent Image Relighting

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:02.009119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:02.009119Z digest=sha256:8534dad212204f0c2c4c184ccd6a986ab674eff2ed6ba7b9be37e3c8b39fc6f3

Observation 6b88b385-05f8-467a-ba60-67e3d5e6c0ee · outbound

This paper cites Zero-shot image harmonization with generative model prior.IEEE Transactions on Multimedia, 27:4494–4507, 2025.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Zero-shot image harmonization with generative model prior.IEEE Transactions on Multimedia, 27:4494–4507, 2025

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:02.217492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:02.217492Z digest=sha256:c23f7ee1fa5f47eab82e4e3373bad17f6a989878795b7beab83c46b1c959b19b

Observation a1eb2549-78f4-42cd-8c3a-5490f1660109 · outbound

This paper cites Barron, Ben Mildenhall, Dor Verbin, Pratul P.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Barron, Ben Mildenhall, Dor Verbin, Pratul P

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:02.394086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:02.394086Z digest=sha256:de8528841a1ab7276a13e2d5efdc6d70e6c74dafce3efc38e467663019e241a6

Observation adfb0d9b-d01f-4611-9d71-cdf930efffca · outbound

This paper cites Nerfstudio: A modular framework for neural radiance field development.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Nerfstudio: A modular framework for neural radiance field development

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:02.560968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:02.560968Z digest=sha256:e0997fc05042084c21e21ae264a618b742a933f79ab037c61690785f9cf03123

Observation 1fc8058e-9be4-4b0a-bc9e-4681cfbfdd9f · outbound

This paper cites 3D Gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics (SIGGRAPH), 2023.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric 3D Gaussian splatting for real-time radiance field rendering.ACM Transactions on Graphics (SIGGRAPH), 2023

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:02.736965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:02.736965Z digest=sha256:62a33fbab522903020ab9814fd3e61afa80f57627016a6755659eb7be60f69c9

Observation c8ed9b99-2309-408a-ac73-96130c155ee7 · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Srinivasan, Matthew Tancik, Jonathan T

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:02.983115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:02.983115Z digest=sha256:6ba670b6840e7259cacd15f083d86268330f8f6aff4fa232e9a8dffc64e1ade0

Observation b1a45cd3-60f6-4210-b20c-e8347c3c10a0 · outbound

This paper cites Instant neural graphics primitives with a multiresolution hash encoding.ACM Transactions on Graphics (SIGGRAPH), 2022.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Instant neural graphics primitives with a multiresolution hash encoding.ACM Transactions on Graphics (SIGGRAPH), 2022

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:03.038997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:03.038997Z digest=sha256:b3b0f6cf4ae2ace20535057f372abd86a92dd86d369d14fcfdfda6b08d22a1a2

Observation e1e6b582-aac2-4fc3-bccb-ff299a674edc · outbound

This paper cites Structured 3d latents for scalable and versatile 3d generation.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Structured 3d latents for scalable and versatile 3d generation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:03.160260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:03.160260Z digest=sha256:bbaa50f77df0c0cd3143e968b5695cb548f18fa73c993a0404cace2dec4f001c

Observation 348631fb-d56d-48f2-9c55-9e47e5dce619 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:03.349347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:03.349347Z digest=sha256:50b884644f5522b8975a5cae11c9c5747609bba856fc2f417709a3290a208e97

Observation 58ec7312-8065-4802-84d5-5e8d46d42aa0 · outbound

This paper cites Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:03.512451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:03.512451Z digest=sha256:cf475cacb148bd648652d1f787bfe089f07fd7c96e49eab5e528b47b6fa04e49

Observation 42a0420c-1712-44b8-a191-a156af54112c · outbound

This paper cites Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:03.699857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:03.699857Z digest=sha256:e586487452b8762b96073d4cbef40a58ed674586a38a7f085bc452e7940873f5

Observation 1c3cfe0f-49e4-46fb-8eac-a26ef5225502 · outbound

This paper cites SGDR: Stochastic gradient descent with warm restarts.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric SGDR: Stochastic gradient descent with warm restarts

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:03.894149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:03.894149Z digest=sha256:d0e3b643f9deda7bdf58e8e0c8dc403ed1e258d9c6d53e4831a905d93b920cc6

Observation ee9e58bf-ba7c-4d0a-a677-e5c9ce041507 · outbound

This paper cites Calibrate before use: Improving few-shot performance of language models.

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Calibrate before use: Improving few-shot performance of language models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-01T15:43:04.068454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:43:04.068454Z digest=sha256:8320484964bd03a3792ade509df8de42f13a88882ce78dbdc576c7872f825b58

Pith citing papers

Observation 238a3b00-e386-4bbc-acf0-02289829a3c9 · inbound

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding cites this paper.

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T11:43:27.836119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-08-05T11:43:26.613972Z digest=sha256:4f79e64f7d070192eef9fb347009616627266159c0eeb9ff78f1b4ddfeea4c27