Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:52:20.853170Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2412.10594.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:52:20.853170Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T00:53:35.188721Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T08:36:26.482689Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 17c76d01-0a43-4413-9ca1-1c11c3c02838 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9ada32-5c20-4329-ae7a-bb6c736e9d9f · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Getting vit in shape: Scaling laws for compute-optimal model design
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fcbf14c3-cec0-4f98-9c87-b6bc37f76105 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Improving image captioning descriptive- ness by ranking and llm-based fusion
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51dfb68-7948-473f-9ac6-d97d834435e5 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Learning a deep single image contrast enhancer from multi-exposure images
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e17d57-450e-4094-9598-a01932b70e4d · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Emerg- ing properties in self-supervised vision transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46faf46c-e022-4296-a3ee-a7738fa3cfaf · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Reproducible scal- ing laws for contrastive language-image learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5301a0e6-3010-4e86-bcfd-0c4e4172cdf5 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Adversarially robust clip mod- els induce better (robust) perceptual metrics
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5f256352-438c-4a75-b6b7-5d347f933941 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Dream- sim: Learning new dimensions of human visual similarity using synthetic data
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 248a84da-d421-42b5-b827-57a691d103b1 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ca037e-9fe7-4594-8b72-ead91024827c · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics R-LPIPS: An adversarially robust perceptual similarity metric
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7553f00d-4b45-46b2-a213-2e46d7755918 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics EMMA: Efficient Visual Alignment in Multi-Modal LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f00ae9a-dee6-4749-a102-8b2fc039c5fc · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Lipsim: A provably robust perceptual similarity metric
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 055ed614-487c-4d54-a188-f7d8395924af · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Generative adversarial networks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f0b984e3-ae3b-41a1-a946-365a82742d77 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Masked autoencoders are scalable vision learners
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 768a2a18-9037-4942-9b8b-e99759ceb529 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Clipscore: A reference-free evaluation met- ric for image captioning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d9e3fd5-4727-4813-983a-4c43c38a8296 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Denoising dif- fusion probabilistic models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4bca766e-af38-420e-b26e-12dec1e77cc4 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Koniq-10k: An ecologically valid database for deep learning of blind image quality assessment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7f102026-893e-49de-817b-2145504c12bb · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics LoRA: Low-rank adaptation of large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0706775c-7b58-4056-854e-575a58f0be4a · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Aesexpert: Towards multi-modality foun- dation model for image aesthetics perception
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed87fb23-beca-42d8-b7e7-1e1386342417 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d7d4595-41c1-44fa-9c73-35d5c952ef8c · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87c04987-a20c-4660-bb21-60d9fa57363f · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Pipal: a large-scale image quality assessment dataset for perceptual image restoration
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6adcabd3-294e-4608-9c12-f8dbd0ee536a · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics ImagenHub: Standardizing the evaluation of conditional image generation models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f42d74-cee5-4d53-8ad1-b622642bf1c1 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ebd9825-6b74-4da0-b548-49e7270fa039 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Agiqa-3k: An open database for ai-generated image quality assessment
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 38f8cb07-b66c-48cc-b277-01fadedc1078 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics AIGIQA-20K: A Large Database for AI-Generated Image Quality Assessment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a2c559-3316-493e-be5f-3d9e2562158b · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75bf0df2-9641-4f84-b2b8-f4038de5a9b9 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6c9e92-3fe4-4e4f-9893-65b40f1222a9 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Kadid-10k: A large-scale artificially distorted iqa database
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2856e819-fd2f-4b73-b33a-94b9572dd042 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Microsoft coco: Common objects in context
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06dafb6f-09b3-47b4-ac97-9522d6245161 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f6262e8d-72fc-4613-89b3-5594fd668549 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Improved baselines with visual instruction tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 290c40b1-d467-494b-8394-25b0777f7ba9 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d45ef75-527d-43ad-a1f5-789ebdbca73f · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e678f20-634d-452a-97a4-631ed9541289 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Vandermeulen, and Simon Kornblith
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0bd7fc4f-a57b-4108-bc20-f44c8c52d153 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Lost in quantization: Improving particu- lar object retrieval in large scale image databases
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6dd01535-2934-4613-a92b-50ebe1af757c · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14160a5e-c550-463e-a728-50a63b09ce36 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Pieapp: Perceptual image-error assessment through pairwise preference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7dc182ff-f0aa-4b51-b3f7-dcbc099e7735 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Learning transferable visual models from natural language supervision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d3e3eb5-774f-4ef3-92a7-aa683fd2a1eb · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Positive-augmented contrastive learning for image and video captioning evaluation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 986fc960-2fd1-42ca-acb7-081ec7faf1c3 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics When Does Perceptual Alignment Benefit Vision Representations?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88bae7ca-aa88-4e61-96b3-2d5eb9dea848 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae66f87-1dc5-4b7f-8485-d15526238d19 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Polos: Multimodal metric learning from human feed- back for image captioning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab87bdbb-d304-417f-b270-810d1ab4a970 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 240ca792-5f11-4660-926e-af9f714084f4 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Ex- ploring clip for assessing the look and feel of images
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6baf2c07-0a34-4d85-b17c-83b77d28dfcb · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53337deb-fa14-454c-bb61-b742d6f0d5e0 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Towards Open-ended Visual Quality Comparison
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83722df9-9853-4a9c-8b18-115afc89969f · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e7f3ee-210f-47cf-a208-2aa100b3880f · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Imagere- ward: Learning and evaluating human preferences for text- to-image generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16d4f2eb-1bb6-4fff-afe1-99534ca33c07 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics mplug- owl2: Revolutionizing multi-modal large language model with modality collaboration
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 021d40ec-759e-4189-9059-8c0083d87534 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Sigmoid Loss for Language Image Pre-Training
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0109bb72-16dd-4d30-afd5-29dd22b2c998 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Text-to-image Diffusion Models in Generative AI: A Survey
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63550c03-bfd8-424b-a306-0f8b6a1c56ca · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Magicbrush: A manually annotated dataset for instruction- guided image editing
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fd1a6379-26b0-4115-86de-fda62eba9f85 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Efros, Eli Shecht- man, and Oliver Wang
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7625820-c4d3-4a4b-808b-23313a459cfb · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Blind image quality assessment via vision- language correspondence: A multitask learning perspective
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation afb4e741-d43f-4779-8027-ba6b89d1f49b · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6155f7f7-7f0b-45a1-b3da-ddc8328ad328 · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics 2AFC Prompting of Large Multimodal Models for Image Quality Assessment
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d03083d-7bac-494f-bcf0-dd6c7656c0fa · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86a91e8-aa96-4272-9426-cdbabaa8de2c · outbound
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 199e40c4-055f-4d48-bd48-e116c72c03bf · inbound
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.