{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:KXRSE3YZOXYB7PQ7CWPXQSWR3Z","short_pith_number":"pith:KXRSE3YZ","schema_version":"1.0","canonical_sha256":"55e3226f1975f01fbe1f159f784ad1de7b2627bc57a7eafc352a89f166a26fdf","source":{"kind":"arxiv","id":"2401.08276","version":1},"attestation_state":"computed","paper":{"title":"AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Haoning Wu, Leida Li, Pengfei Chen, Quan Yuan, Weisi Lin, Xiangfei Sheng, Yipo Huang, Yuzhe Yang, Zhichao Yang","submitted_at":"2024-01-16T10:58:07Z","abstract_excerpt":"With collective endeavors, multimodal large language models (MLLMs) are undergoing a flourishing development. However, their performances on image aesthetics perception remain indeterminate, which is highly desired in real-world applications. An obvious obstacle lies in the absence of a specific benchmark to evaluate the effectiveness of MLLMs on aesthetic perception. This blind groping may impede the further development of more advanced MLLMs with aesthetic perception capacity. To address this dilemma, we propose AesBench, an expert benchmark aiming to comprehensively evaluate the aesthetic p"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2401.08276","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2024-01-16T10:58:07Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"bbe73f0e08f369b1097b13175b59301a0b275a326ade06f51f3a7b69b8189c49","abstract_canon_sha256":"c5b0592368e34e70677bb138b0f252990bd4056f5fcb6390e9ab222e3eb51b04"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:34:09.908665Z","signature_b64":"GgWPnT99/e5KmFfK3Cd5cAK85N92WBU1lL6lAvH0i2dZ3sXShrddvmbHaJe5fcKyZ9rFkWW+HvhI0gN1qmXwDg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"55e3226f1975f01fbe1f159f784ad1de7b2627bc57a7eafc352a89f166a26fdf","last_reissued_at":"2026-07-05T07:34:09.908254Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:34:09.908254Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Haoning Wu, Leida Li, Pengfei Chen, Quan Yuan, Weisi Lin, Xiangfei Sheng, Yipo Huang, Yuzhe Yang, Zhichao Yang","submitted_at":"2024-01-16T10:58:07Z","abstract_excerpt":"With collective endeavors, multimodal large language models (MLLMs) are undergoing a flourishing development. However, their performances on image aesthetics perception remain indeterminate, which is highly desired in real-world applications. An obvious obstacle lies in the absence of a specific benchmark to evaluate the effectiveness of MLLMs on aesthetic perception. This blind groping may impede the further development of more advanced MLLMs with aesthetic perception capacity. To address this dilemma, we propose AesBench, an expert benchmark aiming to comprehensively evaluate the aesthetic p"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2401.08276","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2401.08276/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2401.08276","created_at":"2026-07-05T07:34:09.908320+00:00"},{"alias_kind":"arxiv_version","alias_value":"2401.08276v1","created_at":"2026-07-05T07:34:09.908320+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2401.08276","created_at":"2026-07-05T07:34:09.908320+00:00"},{"alias_kind":"pith_short_12","alias_value":"KXRSE3YZOXYB","created_at":"2026-07-05T07:34:09.908320+00:00"},{"alias_kind":"pith_short_16","alias_value":"KXRSE3YZOXYB7PQ7","created_at":"2026-07-05T07:34:09.908320+00:00"},{"alias_kind":"pith_short_8","alias_value":"KXRSE3YZ","created_at":"2026-07-05T07:34:09.908320+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.05744","citing_title":"PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28696","citing_title":"COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29689","citing_title":"Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2511.16814","citing_title":"Stable diffusion models reveal a persisting human and AI gap in visual creativity","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2601.09896","citing_title":"The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14635","citing_title":"MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2604.12175","citing_title":"Redefining Quality Criteria and Distance-Aware Score Modeling for Image Editing Assessment","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11374","citing_title":"What Do Vision-Language Models Encode for Personalized Image Aesthetics Assessment?","ref_index":3,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z","json":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z.json","graph_json":"https://pith.science/api/pith-number/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/graph.json","events_json":"https://pith.science/api/pith-number/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/events.json","paper":"https://pith.science/paper/KXRSE3YZ"},"agent_actions":{"view_html":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z","download_json":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z.json","view_paper":"https://pith.science/paper/KXRSE3YZ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2401.08276&json=true","fetch_graph":"https://pith.science/api/pith-number/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/graph.json","fetch_events":"https://pith.science/api/pith-number/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/action/timestamp_anchor","attest_storage":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/action/storage_attestation","attest_author":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/action/author_attestation","sign_citation":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/action/citation_signature","submit_replication":"https://pith.science/pith/KXRSE3YZOXYB7PQ7CWPXQSWR3Z/action/replication_record"}},"created_at":"2026-07-05T07:34:09.908320+00:00","updated_at":"2026-07-05T07:34:09.908320+00:00"}