Pith. sign in

Paper Citation Record · LEDGER

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

As of 12 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2412.10594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10594 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:52:20.853170Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T00:53:35.188721Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:36:26.482689Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 17c76d01-0a43-4413-9ca1-1c11c3c02838 · outbound

This paper cites GPT-4 Technical Report.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.513627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.513627Z digest=sha256:55b370e53ac63fa666e1d6a510314dea0401e3f565581ecc23f081d0e3752551

Observation 8a9ada32-5c20-4329-ae7a-bb6c736e9d9f · outbound

This paper cites Getting vit in shape: Scaling laws for compute-optimal model design.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Getting vit in shape: Scaling laws for compute-optimal model design

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:22.028682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.519143Z digest=sha256:0a2704d9df28d6459059762bd1fae13b5c8368d1c9f6f70ea5301c6b1ec8a5fc

Observation fcbf14c3-cec0-4f98-9c87-b6bc37f76105 · outbound

This paper cites Improving image captioning descriptive- ness by ranking and llm-based fusion.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Improving image captioning descriptive- ness by ranking and llm-based fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.524263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.524263Z digest=sha256:69c65113871561216a7554e631108613737a495a57d2ef5cca468262a862f842

Observation c51dfb68-7948-473f-9ac6-d97d834435e5 · outbound

This paper cites Learning a deep single image contrast enhancer from multi-exposure images.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Learning a deep single image contrast enhancer from multi-exposure images

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.528675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.528675Z digest=sha256:d51322ce9d3876f847a95537ac23df703beeb99e77402a883f099e6f954f5951

Observation 77e17d57-450e-4094-9598-a01932b70e4d · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.991793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.532907Z digest=sha256:18efdeea6a752bbf1ad8a1952aa290dce66e0ac05cdc62f5a071b1a5805aeb38

Observation 46faf46c-e022-4296-a3ee-a7738fa3cfaf · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.537679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.537679Z digest=sha256:890cf788c6f3df3229cbe507665c6fcac4880e80354099194b1d8a444b317d27

Observation 5301a0e6-3010-4e86-bcfd-0c4e4172cdf5 · outbound

This paper cites Adversarially robust clip mod- els induce better (robust) perceptual metrics.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Adversarially robust clip mod- els induce better (robust) perceptual metrics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.957738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.544410Z digest=sha256:521807e63f5491c3cb4afc5141c7f702306bcb44f48c6640a474713e1bf6ff2c

Observation 5f256352-438c-4a75-b6b7-5d347f933941 · outbound

This paper cites Dream- sim: Learning new dimensions of human visual similarity using synthetic data.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Dream- sim: Learning new dimensions of human visual similarity using synthetic data

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.936354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.551039Z digest=sha256:1e0023cf795cd18ad0029783eff063bf091896b4ac77e640037bbad425d4a7dd

Observation 248a84da-d421-42b5-b827-57a691d103b1 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.559022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.559022Z digest=sha256:aa2165c92eb6a8c8391bcb4cc0160a906a96a04f996aa4185fa1dd1fc5d25067

Observation 78ca037e-9fe7-4594-8b72-ead91024827c · outbound

This paper cites R-LPIPS: An adversarially robust perceptual similarity metric.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics R-LPIPS: An adversarially robust perceptual similarity metric

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.921279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.564908Z digest=sha256:5a777cf8276c25e0f2c1063de717b83ddb94ab7e5230ec25a4af97014b902d2f

Observation 7553f00d-4b45-46b2-a213-2e46d7755918 · outbound

This paper cites EMMA: Efficient Visual Alignment in Multi-Modal LLMs.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics EMMA: Efficient Visual Alignment in Multi-Modal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.570468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.570468Z digest=sha256:8993ab40107805d60d71956bdbc99888293c58d7f9a7d380b2bbadc18091ddeb

Observation 2f00ae9a-dee6-4749-a102-8b2fc039c5fc · outbound

This paper cites Lipsim: A provably robust perceptual similarity metric.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Lipsim: A provably robust perceptual similarity metric

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.904307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.575211Z digest=sha256:9468e65cf552aa21c7202d1ec20d51ca29f7ef09c193ba5b5abbc53c4ab8a3bb

Observation 055ed614-487c-4d54-a188-f7d8395924af · outbound

This paper cites Generative adversarial networks.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Generative adversarial networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.888020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.579289Z digest=sha256:8356388164601d0cb32fe33f55f28ee7d816cb94c4a3e1c2e10183efeb4459ce

Observation f0b984e3-ae3b-41a1-a946-365a82742d77 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Masked autoencoders are scalable vision learners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.874994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.583662Z digest=sha256:75ffe3e86d22647e8c650d4b5f0e9cd783e6a77b28726da67afccea21ac3b2ef

Observation 768a2a18-9037-4942-9b8b-e99759ceb529 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Clipscore: A reference-free evaluation met- ric for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.861345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.588191Z digest=sha256:1663743cd1322ce474a9d103b9697ce367855148fab96b2e07c1771fddd745ce

Observation 7d9e3fd5-4727-4813-983a-4c43c38a8296 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Denoising dif- fusion probabilistic models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.845892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.593566Z digest=sha256:a513d8a5120fc86c4c46672a7aedefa5c7994678371632f95ead2896f6dc3101

Observation 4bca766e-af38-420e-b26e-12dec1e77cc4 · outbound

This paper cites Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.829191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.598218Z digest=sha256:0e0d5f857010fcbc66f4b4c49cbdd615d2da87406414fbd16fa9e0f10aa20334

Observation 7f102026-893e-49de-817b-2145504c12bb · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics LoRA: Low-rank adaptation of large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.812630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.602776Z digest=sha256:048398e1b3a0d977a7a563a410383d41390a820f8d646c0868becb3ba21b16f9

Observation 0706775c-7b58-4056-854e-575a58f0be4a · outbound

This paper cites Aesexpert: Towards multi-modality foun- dation model for image aesthetics perception.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Aesexpert: Towards multi-modality foun- dation model for image aesthetics perception

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.797672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.607404Z digest=sha256:614a489a7e736f37cfdafb848f4ce8f53605c82a4d51d823aea3f6a2db0a3322

Observation ed87fb23-beca-42d8-b7e7-1e1386342417 · outbound

This paper cites HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.612245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.612245Z digest=sha256:50a6a9fc0ee73c96f17e916d96df0defb23d8b19a39fcdefc47775d70cfd1e7c

Observation 2d7d4595-41c1-44fa-9c73-35d5c952ef8c · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.618332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.618332Z digest=sha256:ab9a16c11d1e4bf0083a833786f5894cbeda71dc538e880ccfa74d44679f730e

Observation 87c04987-a20c-4660-bb21-60d9fa57363f · outbound

This paper cites Pipal: a large-scale image quality assessment dataset for perceptual image restoration.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Pipal: a large-scale image quality assessment dataset for perceptual image restoration

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.782689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.623290Z digest=sha256:9c04e7303630f6ecd497d6a4a0d44ab4f931730917330a289d2de696aadeb62b

Observation 6adcabd3-294e-4608-9c12-f8dbd0ee536a · outbound

This paper cites ImagenHub: Standardizing the evaluation of conditional image generation models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics ImagenHub: Standardizing the evaluation of conditional image generation models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.628436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.628436Z digest=sha256:f5201990f1863a9b954ee10b694c29d329cc15bda4eab1fba63561ae88eb0d44

Observation 87f42d74-cee5-4d53-8ad1-b622642bf1c1 · outbound

This paper cites UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.635404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.635404Z digest=sha256:cccd91386e184aee6d28f0b59629af2d469971be7595a9849f231467cb28a14d

Observation 0ebd9825-6b74-4da0-b548-49e7270fa039 · outbound

This paper cites Agiqa-3k: An open database for ai-generated image quality assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Agiqa-3k: An open database for ai-generated image quality assessment

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.766045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.640156Z digest=sha256:0b0da00ba58274009703b7f3f3a109f36110cc17c4bb542aeb47bbe92ba2a893

Observation 38f8cb07-b66c-48cc-b277-01fadedc1078 · outbound

This paper cites AIGIQA-20K: A Large Database for AI-Generated Image Quality Assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics AIGIQA-20K: A Large Database for AI-Generated Image Quality Assessment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.645204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.645204Z digest=sha256:25539599247ca93699dd5877a6c5035552269b865d0265819c69325d87abea08

Observation 15a2c559-3316-493e-be5f-3d9e2562158b · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.651720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.651720Z digest=sha256:3cbbb71233266812c326ec1962633d388f5df915a6e235bc96bc1de1257f89ea

Observation 75bf0df2-9641-4f84-b2b8-f4038de5a9b9 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.657956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.657956Z digest=sha256:e371d2b7fd8c74d89309d74484ce577dea5885bd88989c15cb0b1e9f9ab6501d

Observation 8b6c9e92-3fe4-4e4f-9893-65b40f1222a9 · outbound

This paper cites Kadid-10k: A large-scale artificially distorted iqa database.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Kadid-10k: A large-scale artificially distorted iqa database

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.747189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.663794Z digest=sha256:00329f481b690f87bffa339299f9d69a7d12f698b29927f87b07606c27ba5de7

Observation 2856e819-fd2f-4b73-b33a-94b9572dd042 · outbound

This paper cites Microsoft coco: Common objects in context.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Microsoft coco: Common objects in context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.731632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.668842Z digest=sha256:bb80a7270e6b943b074f9ce942c20c84cefe6864c174d811e98ded2cdcc8b1de

Observation 06dafb6f-09b3-47b4-ac97-9522d6245161 · outbound

This paper cites an unresolved cited work.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:52:21.713584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.674338Z digest=sha256:4c8135cb2949193c6e49d516f2adb43678b6761355e642c3c4f3ab58a22e17ca

Observation f6262e8d-72fc-4613-89b3-5594fd668549 · outbound

This paper cites Improved baselines with visual instruction tuning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Improved baselines with visual instruction tuning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.699013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.679260Z digest=sha256:a74f3601846bfafdc4d6b5a758fa2de056715d3a534201120164a21138610e60

Observation 290c40b1-d467-494b-8394-25b0777f7ba9 · outbound

This paper cites FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.683958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.683958Z digest=sha256:b28b4edcd3afa929986ed861913e9221aa99791607026444efc3ea13c1cdfffd

Observation 6d45ef75-527d-43ad-a1f5-789ebdbca73f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.689029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.689029Z digest=sha256:656f3e27627433802e2485e0b4e0f300a7310be49348b297b871a95eaff9e253

Observation 1e678f20-634d-452a-97a4-631ed9541289 · outbound

This paper cites Vandermeulen, and Simon Kornblith.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Vandermeulen, and Simon Kornblith

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.680734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.694008Z digest=sha256:0877d484042741b16cdec32ae2ab630420d2bf20c5c9a7426932e8d7ec7ec2a8

Observation 0bd7fc4f-a57b-4108-bc20-f44c8c52d153 · outbound

This paper cites Lost in quantization: Improving particu- lar object retrieval in large scale image databases.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Lost in quantization: Improving particu- lar object retrieval in large scale image databases

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.665683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.699753Z digest=sha256:612274773348aa3f463039ca90fa20d565f1d689c49fdb7c59f94f8cd9591345

Observation 6dd01535-2934-4613-a92b-50ebe1af757c · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.706284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.706284Z digest=sha256:8490acf2ea4b7b7d2d19d65d4289740f51f79be56f874ee5072a05af21675ca1

Observation 14160a5e-c550-463e-a728-50a63b09ce36 · outbound

This paper cites Pieapp: Perceptual image-error assessment through pairwise preference.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Pieapp: Perceptual image-error assessment through pairwise preference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.649485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.714728Z digest=sha256:f290bef30e2c2a4aa4476f0a0ebdd9e475fcf8127b69207b0e0006e176237e27

Observation 7dc182ff-f0aa-4b51-b3f7-dcbc099e7735 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Learning transferable visual models from natural language supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.634785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.719641Z digest=sha256:c443fa3c39a3c242519ac51e61352a9bc374df3d741307158bb8d48935a6f964

Observation 6d3e3eb5-774f-4ef3-92a7-aa683fd2a1eb · outbound

This paper cites Positive-augmented contrastive learning for image and video captioning evaluation.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Positive-augmented contrastive learning for image and video captioning evaluation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.619423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.725645Z digest=sha256:fe7f19b734f0884f83b703f070cf0bc13f718b08aef3efa00d4699816e29f2bc

Observation 986fc960-2fd1-42ca-acb7-081ec7faf1c3 · outbound

This paper cites When Does Perceptual Alignment Benefit Vision Representations?.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics When Does Perceptual Alignment Benefit Vision Representations?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.730292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.730292Z digest=sha256:ebcc7409b43693370c7bd47fb7fa5d10f2f5c7f66d37947da1797dd5e9cf1158

Observation 88bae7ca-aa88-4e61-96b3-2d5eb9dea848 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.735706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.735706Z digest=sha256:fb2a4a5e1efaa07f17998f62229d645b247355234d6697c38fd61bf19079198c

Observation 9ae66f87-1dc5-4b7f-8485-d15526238d19 · outbound

This paper cites Polos: Multimodal metric learning from human feed- back for image captioning.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Polos: Multimodal metric learning from human feed- back for image captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.604954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.741709Z digest=sha256:835462a4c24f2e70ede02085987b737adc4fefa8f7feab50445a19435a977584

Observation ab87bdbb-d304-417f-b270-810d1ab4a970 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.746009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.746009Z digest=sha256:431e0cb96c2963909df693f2123c387ace662e3ed78e766cdfe63bd65f67fbb1

Observation 240ca792-5f11-4660-926e-af9f714084f4 · outbound

This paper cites Ex- ploring clip for assessing the look and feel of images.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Ex- ploring clip for assessing the look and feel of images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.588807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.751054Z digest=sha256:b2134534d3af8e5f331664c8164c640d20b6e0e42edd047a161e537a3ab722af

Observation 6baf2c07-0a34-4d85-b17c-83b77d28dfcb · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.757343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.757343Z digest=sha256:7a202c9cc7f723697b28d095399148272704e954e5af1ac3d468e93776ebf8ee

Observation 53337deb-fa14-454c-bb61-b742d6f0d5e0 · outbound

This paper cites Towards Open-ended Visual Quality Comparison.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Towards Open-ended Visual Quality Comparison

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.764418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.764418Z digest=sha256:eaee3c6d491f4aa5c1e3c1aa8badfdd610fea409d00ee8bb7dd7a33ba9af40a2

Observation 83722df9-9853-4a9c-8b18-115afc89969f · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.771725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.771725Z digest=sha256:ea5fc9ce71baab691a564f44604072b41801a847a4acb738889d2a356da44aa7

Observation 75e7f3ee-210f-47cf-a208-2aa100b3880f · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.546559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.783050Z digest=sha256:ab8d951f991287ad4cbdf3d321de169f4ca901b19f7f114c0c99a5fadea0f3c2

Observation 16d4f2eb-1bb6-4fff-afe1-99534ca33c07 · outbound

This paper cites mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.530751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.790872Z digest=sha256:f0a17d67e8de26098824475b5e95a5d46cbcd68797c336f29f2bff5e476e83ad

Observation 021d40ec-759e-4189-9059-8c0083d87534 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Sigmoid Loss for Language Image Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.798079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.798079Z digest=sha256:a84d42712f9aef0efd4a6d1bc48f3b68b459ec8c3d6e684bd2d8578d90c16f6b

Observation 0109bb72-16dd-4d30-afd5-29dd22b2c998 · outbound

This paper cites Text-to-image Diffusion Models in Generative AI: A Survey.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Text-to-image Diffusion Models in Generative AI: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.803758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.803758Z digest=sha256:8fb636893ca5e383e21e2bd588b2a53ebce1efc39ce0815c8dfa28006dec63b4

Observation 63550c03-bfd8-424b-a306-0f8b6a1c56ca · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.513857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.812589Z digest=sha256:983c7e67600833b6730bbdb8b2b3354278cd1baa92e9833d7dba6cd5eb3fc4cb

Observation fd1a6379-26b0-4115-86de-fda62eba9f85 · outbound

This paper cites Efros, Eli Shecht- man, and Oliver Wang.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Efros, Eli Shecht- man, and Oliver Wang

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.496296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.823969Z digest=sha256:f8267d90cdbbcac422cce9d322a89f94a5aba7b7ace6709ff2cd60bbd34f86b0

Observation c7625820-c4d3-4a4b-808b-23313a459cfb · outbound

This paper cites Blind image quality assessment via vision- language correspondence: A multitask learning perspective.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Blind image quality assessment via vision- language correspondence: A multitask learning perspective

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:52:21.480909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.831283Z digest=sha256:ea4d7d3f5947829db773ac3823a3c5f3ef7370228085332f07477aa16c6192dc

Observation afb4e741-d43f-4779-8027-ba6b89d1f49b · outbound

This paper cites A-Bench: Are LMMs Masters at Evaluating AI-generated Images?.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics A-Bench: Are LMMs Masters at Evaluating AI-generated Images?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.839190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.839190Z digest=sha256:4a77816a92c365a4f128aa720f3f3696752a04c74888810c62a1116bae146b94

Observation 6155f7f7-7f0b-45a1-b3da-ddc8328ad328 · outbound

This paper cites 2AFC Prompting of Large Multimodal Models for Image Quality Assessment.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics 2AFC Prompting of Large Multimodal Models for Image Quality Assessment

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.845040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.845040Z digest=sha256:f8a4140a2391aeb3f0f7e5ba52a6041ddb4bc2a595558009eb520340a6e9262f

Observation 8d03083d-7bac-494f-bcf0-dd6c7656c0fa · outbound

This paper cites Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:52:20.853170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:52:20.853170Z digest=sha256:47f572a8e6715a71ae6156a8e2a12cd02ca729d870b31cf4f8e36f82e018785d

Observation d86a91e8-aa96-4272-9426-cdbabaa8de2c · outbound

This paper cites an unresolved cited work.

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:52:21.563678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:52:20.777233Z digest=sha256:84ab644e29f53094dd13e5b80e77a34b3247f27aa92315dc1f7fb4b9f231673b

Pith citing papers

Observation 199e40c4-055f-4d48-bd48-e116c72c03bf · inbound

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding cites this paper.

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.485902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T00:53:35.188721Z digest=sha256:ba19a744c7768204a09ac48e09a9f7f42edbb418c13be7ab1b838a2a2eb193d7