{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:4P7EJVN5HMABSQJLDRWT3KKA4E","short_pith_number":"pith:4P7EJVN5","schema_version":"1.0","canonical_sha256":"e3fe44d5bd3b0019412b1c6d3da940e1372c5615b9f1318a115dfaaf155482e8","source":{"kind":"arxiv","id":"2311.03287","version":2},"attestation_state":"computed","paper":{"title":"Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL","cs.CV"],"primary_cat":"cs.LG","authors_text":"Chenhang Cui, Huaxiu Yao, James Zou, Linjun Zhang, Shirley Wu, Xinyu Yang, Yiyang Zhou","submitted_at":"2023-11-06T17:26:59Z","abstract_excerpt":"While GPT-4V(ision) impressively models both visual and textual information simultaneously, it's hallucination behavior has not been systematically assessed. To bridge this gap, we introduce a new benchmark, namely, the Bias and Interference Challenges in Visual Language Models (Bingo). This benchmark is designed to evaluate and shed light on the two common types of hallucinations in visual language models: bias and interference. Here, bias refers to the model's tendency to hallucinate certain types of responses, possibly due to imbalance in its training data. Interference pertains to scenario"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2311.03287","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2023-11-06T17:26:59Z","cross_cats_sorted":["cs.CL","cs.CV"],"title_canon_sha256":"961c8153e8fb898cf0c24249d52a39b5c69956fd788de0de0fb439b7dd7e7e05","abstract_canon_sha256":"d8a39b6873aff4cc478b0b7e0a4a207a3fe1b2a63a8ac00b4099df07c356b7a7"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:09:56.677336Z","signature_b64":"HMbUmLaGtGC2fk6Ql1Rx7abEa502k/qZ3ckvGQBOr5ro7X5E7MpGyT5AWbASPmFzgu5NjbKx09+3YO7GuKTzAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e3fe44d5bd3b0019412b1c6d3da940e1372c5615b9f1318a115dfaaf155482e8","last_reissued_at":"2026-07-05T07:09:56.676825Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:09:56.676825Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL","cs.CV"],"primary_cat":"cs.LG","authors_text":"Chenhang Cui, Huaxiu Yao, James Zou, Linjun Zhang, Shirley Wu, Xinyu Yang, Yiyang Zhou","submitted_at":"2023-11-06T17:26:59Z","abstract_excerpt":"While GPT-4V(ision) impressively models both visual and textual information simultaneously, it's hallucination behavior has not been systematically assessed. To bridge this gap, we introduce a new benchmark, namely, the Bias and Interference Challenges in Visual Language Models (Bingo). This benchmark is designed to evaluate and shed light on the two common types of hallucinations in visual language models: bias and interference. Here, bias refers to the model's tendency to hallucinate certain types of responses, possibly due to imbalance in its training data. Interference pertains to scenario"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2311.03287","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2311.03287/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2311.03287","created_at":"2026-07-05T07:09:56.676887+00:00"},{"alias_kind":"arxiv_version","alias_value":"2311.03287v2","created_at":"2026-07-05T07:09:56.676887+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2311.03287","created_at":"2026-07-05T07:09:56.676887+00:00"},{"alias_kind":"pith_short_12","alias_value":"4P7EJVN5HMAB","created_at":"2026-07-05T07:09:56.676887+00:00"},{"alias_kind":"pith_short_16","alias_value":"4P7EJVN5HMABSQJL","created_at":"2026-07-05T07:09:56.676887+00:00"},{"alias_kind":"pith_short_8","alias_value":"4P7EJVN5","created_at":"2026-07-05T07:09:56.676887+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":16,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2607.01982","citing_title":"MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30911","citing_title":"What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2411.15594","citing_title":"A Survey on LLM-as-a-Judge","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2507.13868","citing_title":"When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2511.10287","citing_title":"OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2402.11411","citing_title":"Aligning Modalities in Vision Large Language Models via Preference Fine-tuning","ref_index":148,"is_internal_anchor":false},{"citing_arxiv_id":"2310.14566","citing_title":"HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2506.20670","citing_title":"MMSearch-R1: Incentivizing LMMs to Search","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2311.07397","citing_title":"AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2311.16502","citing_title":"MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2409.02813","citing_title":"MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09429","citing_title":"Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20696","citing_title":"R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2404.18930","citing_title":"Hallucination of Multimodal Large Language Models: A Survey","ref_index":38,"is_internal_anchor":false},{"citing_arxiv_id":"2604.10219","citing_title":"Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models","ref_index":90,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18803","citing_title":"LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models","ref_index":7,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E","json":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E.json","graph_json":"https://pith.science/api/pith-number/4P7EJVN5HMABSQJLDRWT3KKA4E/graph.json","events_json":"https://pith.science/api/pith-number/4P7EJVN5HMABSQJLDRWT3KKA4E/events.json","paper":"https://pith.science/paper/4P7EJVN5"},"agent_actions":{"view_html":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E","download_json":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E.json","view_paper":"https://pith.science/paper/4P7EJVN5","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2311.03287&json=true","fetch_graph":"https://pith.science/api/pith-number/4P7EJVN5HMABSQJLDRWT3KKA4E/graph.json","fetch_events":"https://pith.science/api/pith-number/4P7EJVN5HMABSQJLDRWT3KKA4E/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E/action/timestamp_anchor","attest_storage":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E/action/storage_attestation","attest_author":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E/action/author_attestation","sign_citation":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E/action/citation_signature","submit_replication":"https://pith.science/pith/4P7EJVN5HMABSQJLDRWT3KKA4E/action/replication_record"}},"created_at":"2026-07-05T07:09:56.676887+00:00","updated_at":"2026-07-05T07:09:56.676887+00:00"}