{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:4QMBPWEV753MSZFH4YSJGD4HVR","short_pith_number":"pith:4QMBPWEV","schema_version":"1.0","canonical_sha256":"e41817d895ff76c964a7e624930f87ac49acc6a189d2e4522a183784060a9941","source":{"kind":"arxiv","id":"2311.07397","version":2},"attestation_state":"computed","paper":{"title":"AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"AMBER provides an LLM-free benchmark to evaluate hallucinations in multi-modal models across existence, attribute and relation dimensions for generative and discriminative tasks.","cross_cats":["cs.CV"],"primary_cat":"cs.CL","authors_text":"Guohai Xu, Haitao Jia, Haiyang Xu, Jiaqi Wang, Jing Zhang, Jitao Sang, Ji Zhang, Junyang Wang, Ming Yan, Yuhang Wang, Yukai Gu","submitted_at":"2023-11-13T15:25:42Z","abstract_excerpt":"Despite making significant progress in multi-modal tasks, current Multi-modal Large Language Models (MLLMs) encounter the significant challenge of hallucinations, which may lead to harmful consequences. Therefore, evaluating MLLMs' hallucinations is becoming increasingly important in model improvement and practical application deployment. Previous works are limited in high evaluation costs (e.g., relying on humans or advanced LLMs) and insufficient evaluation dimensions (e.g., types of tasks and hallucinations). In this paper, we propose an LLM-free multi-dimensional benchmark AMBER, which can"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2311.07397","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2023-11-13T15:25:42Z","cross_cats_sorted":["cs.CV"],"title_canon_sha256":"86a0484be9d98694bb255b98a4a8cc2ddd51a86237d4d5c0e8d1ab7780b830ad","abstract_canon_sha256":"d3b64a4984882eb45d3a0c8ed5040f1b03a382944439c0dc3cc149470cabe60e"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-05-17T23:38:48.795287Z","signature_b64":"TE5d6n7VZOnWJxNERLnDp4JuXzw9P+HVy4leUR8mhK3qjghGz62DMD8YIIBotGxOLogKzh8HwgXwsPyV70yfDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e41817d895ff76c964a7e624930f87ac49acc6a189d2e4522a183784060a9941","last_reissued_at":"2026-05-17T23:38:48.794685Z","signature_status":"signed_v1","first_computed_at":"2026-05-17T23:38:48.794685Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"AMBER provides an LLM-free benchmark to evaluate hallucinations in multi-modal models across existence, attribute and relation dimensions for generative and discriminative tasks.","cross_cats":["cs.CV"],"primary_cat":"cs.CL","authors_text":"Guohai Xu, Haitao Jia, Haiyang Xu, Jiaqi Wang, Jing Zhang, Jitao Sang, Ji Zhang, Junyang Wang, Ming Yan, Yuhang Wang, Yukai Gu","submitted_at":"2023-11-13T15:25:42Z","abstract_excerpt":"Despite making significant progress in multi-modal tasks, current Multi-modal Large Language Models (MLLMs) encounter the significant challenge of hallucinations, which may lead to harmful consequences. Therefore, evaluating MLLMs' hallucinations is becoming increasingly important in model improvement and practical application deployment. Previous works are limited in high evaluation costs (e.g., relying on humans or advanced LLMs) and insufficient evaluation dimensions (e.g., types of tasks and hallucinations). In this paper, we propose an LLM-free multi-dimensional benchmark AMBER, which can"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"we propose an LLM-free multi-dimensional benchmark AMBER, which can be used to evaluate both generative task and discriminative task including existence, attribute and relation hallucination.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the proposed low-cost evaluation pipeline can accurately detect and categorize hallucinations without introducing new biases or missing important cases that would require LLM or human judgment.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"AMBER is an LLM-free multi-dimensional benchmark for evaluating hallucinations in MLLMs across generative and discriminative tasks.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"AMBER provides an LLM-free benchmark to evaluate hallucinations in multi-modal models across existence, attribute and relation dimensions for generative and discriminative tasks.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"f49616b847b32a83bd49ca8fe67a92c3701ae0dc7720b559193bd44db417d7a6"},"source":{"id":"2311.07397","kind":"arxiv","version":2},"verdict":{"id":"a871dcbe-0595-4a3b-b458-3b53fe7ccfc7","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-16T06:41:41.603498Z","strongest_claim":"we propose an LLM-free multi-dimensional benchmark AMBER, which can be used to evaluate both generative task and discriminative task including existence, attribute and relation hallucination.","one_line_summary":"AMBER is an LLM-free multi-dimensional benchmark for evaluating hallucinations in MLLMs across generative and discriminative tasks.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the proposed low-cost evaluation pipeline can accurately detect and categorize hallucinations without introducing new biases or missing important cases that would require LLM or human judgment.","pith_extraction_headline":"AMBER provides an LLM-free benchmark to evaluate hallucinations in multi-modal models across existence, attribute and relation dimensions for generative and discriminative tasks."},"references":{"count":15,"sample":[{"doi":"","year":null,"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","ref_index":1,"cited_arxiv_id":"2308.12966","is_internal_anchor":true},{"doi":"","year":null,"title":"MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning","work_id":"fb62cd1b-3991-40be-a987-3cfa5772b5b5","ref_index":2,"cited_arxiv_id":"2310.09478","is_internal_anchor":true},{"doi":"","year":null,"title":"Holistic analysis of hallucination in gpt-4v(ision): Bias and interference challenges.CoRR, abs/2311.03287","work_id":"a39b177c-6624-4310-9837-526645915677","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning","work_id":"f3aac728-ded0-4e55-aa9e-4a1635d4313d","ref_index":4,"cited_arxiv_id":"2305.06500","is_internal_anchor":true},{"doi":"","year":null,"title":"Detecting and preventing hallucinations in large vi- sion language models","work_id":"0a8f3afc-fe73-4c85-9d6e-217568430f4f","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":15,"snapshot_sha256":"cb4200dcdfc4007c502e276369067b96c8a65bde3540c49f1cda50423e8c712b","internal_anchors":8},"formal_canon":{"evidence_count":1,"snapshot_sha256":"163f2e371f43af8df3c1f1ab64fb30c33d25638150937fce350fe2024eefac07"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2311.07397","created_at":"2026-05-17T23:38:48.794772+00:00"},{"alias_kind":"arxiv_version","alias_value":"2311.07397v2","created_at":"2026-05-17T23:38:48.794772+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2311.07397","created_at":"2026-05-17T23:38:48.794772+00:00"},{"alias_kind":"pith_short_12","alias_value":"4QMBPWEV753M","created_at":"2026-05-18T12:33:33.725879+00:00"},{"alias_kind":"pith_short_16","alias_value":"4QMBPWEV753MSZFH","created_at":"2026-05-18T12:33:33.725879+00:00"},{"alias_kind":"pith_short_8","alias_value":"4QMBPWEV","created_at":"2026-05-18T12:33:33.725879+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":31,"internal_anchor_count":31,"sample":[{"citing_arxiv_id":"2411.16771","citing_title":"VidHal: Benchmarking Temporal Hallucinations in Vision LLMs","ref_index":52,"is_internal_anchor":true},{"citing_arxiv_id":"2509.15435","citing_title":"ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2511.14159","citing_title":"MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs","ref_index":54,"is_internal_anchor":true},{"citing_arxiv_id":"2605.20448","citing_title":"Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?","ref_index":31,"is_internal_anchor":true},{"citing_arxiv_id":"2605.16953","citing_title":"How do Humans Process AI-generated Hallucination Contents: a Neuroimaging Study","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2605.15300","citing_title":"Deep Pre-Alignment for VLMs","ref_index":79,"is_internal_anchor":true},{"citing_arxiv_id":"2505.21472","citing_title":"Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2506.13130","citing_title":"ZINA: Multimodal Fine-grained Hallucination Detection and Editing","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2506.09522","citing_title":"Revisit What You See: Revealing Visual Semantics in Vision Tokens to Guide LVLM Decoding","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2507.21584","citing_title":"TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2509.15435","citing_title":"ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2510.21122","citing_title":"NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation","ref_index":50,"is_internal_anchor":true},{"citing_arxiv_id":"2511.20032","citing_title":"Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2306.13549","citing_title":"A Survey on Multimodal Large Language Models","ref_index":165,"is_internal_anchor":true},{"citing_arxiv_id":"2605.14621","citing_title":"Do We Really Need External Tools to Mitigate Hallucinations? SIRA: Shared-Prefix Internal Reconstruction of Attribution","ref_index":36,"is_internal_anchor":true},{"citing_arxiv_id":"2605.13080","citing_title":"Learning to See What You Need: Gaze Attention for Multimodal Large Language Models","ref_index":107,"is_internal_anchor":true},{"citing_arxiv_id":"2605.13156","citing_title":"Dual-Pathway Circuits of Object Hallucination in Vision-Language Models","ref_index":31,"is_internal_anchor":true},{"citing_arxiv_id":"2605.13173","citing_title":"OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems","ref_index":31,"is_internal_anchor":true},{"citing_arxiv_id":"2402.00253","citing_title":"A Survey on Hallucination in Large Vision-Language Models","ref_index":44,"is_internal_anchor":true},{"citing_arxiv_id":"2605.09443","citing_title":"Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent","ref_index":27,"is_internal_anchor":true},{"citing_arxiv_id":"2605.10622","citing_title":"Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination","ref_index":122,"is_internal_anchor":true},{"citing_arxiv_id":"2605.04874","citing_title":"Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models","ref_index":33,"is_internal_anchor":true},{"citing_arxiv_id":"2604.20696","citing_title":"R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs","ref_index":46,"is_internal_anchor":true},{"citing_arxiv_id":"2404.18930","citing_title":"Hallucination of Multimodal Large Language Models: A Survey","ref_index":164,"is_internal_anchor":true},{"citing_arxiv_id":"2604.07914","citing_title":"Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction","ref_index":44,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":1,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR","json":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR.json","graph_json":"https://pith.science/api/pith-number/4QMBPWEV753MSZFH4YSJGD4HVR/graph.json","events_json":"https://pith.science/api/pith-number/4QMBPWEV753MSZFH4YSJGD4HVR/events.json","paper":"https://pith.science/paper/4QMBPWEV"},"agent_actions":{"view_html":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR","download_json":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR.json","view_paper":"https://pith.science/paper/4QMBPWEV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2311.07397&json=true","fetch_graph":"https://pith.science/api/pith-number/4QMBPWEV753MSZFH4YSJGD4HVR/graph.json","fetch_events":"https://pith.science/api/pith-number/4QMBPWEV753MSZFH4YSJGD4HVR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR/action/storage_attestation","attest_author":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR/action/author_attestation","sign_citation":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR/action/citation_signature","submit_replication":"https://pith.science/pith/4QMBPWEV753MSZFH4YSJGD4HVR/action/replication_record"}},"created_at":"2026-05-17T23:38:48.794772+00:00","updated_at":"2026-05-17T23:38:48.794772+00:00"}