{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:555TZ7WGIQESUA54MQKZTEW5X4","short_pith_number":"pith:555TZ7WG","schema_version":"1.0","canonical_sha256":"ef7b3cfec644092a03bc64159992ddbf37951e3046db8bc43f435e7caacee1ad","source":{"kind":"arxiv","id":"2305.10355","version":3},"attestation_state":"computed","paper":{"title":"Evaluating Object Hallucination in Large Vision-Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Large vision-language models often describe objects absent from the given image, especially those frequent in instructions or co-occurring with visible items.","cross_cats":["cs.CL","cs.MM"],"primary_cat":"cs.CV","authors_text":"Jinpeng Wang, Ji-Rong Wen, Kun Zhou, Wayne Xin Zhao, Yifan Du, Yifan Li","submitted_at":"2023-05-17T16:34:01Z","abstract_excerpt":"Inspired by the superior language abilities of large language models (LLM), large vision-language models (LVLM) have been recently explored by integrating powerful LLMs for improving the performance on complex multimodal tasks. Despite the promising progress on LVLMs, we find that LVLMs suffer from the hallucination problem, i.e. they tend to generate objects that are inconsistent with the target images in the descriptions. To investigate it, this work presents the first systematic study on object hallucination of LVLMs. We conduct the evaluation experiments on several representative LVLMs, an"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2305.10355","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2023-05-17T16:34:01Z","cross_cats_sorted":["cs.CL","cs.MM"],"title_canon_sha256":"58ea4de77f6f74bc1927472cc5d620ed7d1e0d77b59e3fd907bb8cb07f8d5f53","abstract_canon_sha256":"931ab8ddd1af19b56486cde173c84c852ca8f723971792b5bbdde6d56726f174"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:05:08.636981Z","signature_b64":"RdQ0B88jsaKGf63p7VOU+XmILCnvfyUI2TgR9CWzh3/qwg6ELfzyBbzqulZWOOkm2puc4A6vYVfM7BvCh6xNCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"ef7b3cfec644092a03bc64159992ddbf37951e3046db8bc43f435e7caacee1ad","last_reissued_at":"2026-07-05T07:05:08.636573Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:05:08.636573Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Evaluating Object Hallucination in Large Vision-Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Large vision-language models often describe objects absent from the given image, especially those frequent in instructions or co-occurring with visible items.","cross_cats":["cs.CL","cs.MM"],"primary_cat":"cs.CV","authors_text":"Jinpeng Wang, Ji-Rong Wen, Kun Zhou, Wayne Xin Zhao, Yifan Du, Yifan Li","submitted_at":"2023-05-17T16:34:01Z","abstract_excerpt":"Inspired by the superior language abilities of large language models (LLM), large vision-language models (LVLM) have been recently explored by integrating powerful LLMs for improving the performance on complex multimodal tasks. Despite the promising progress on LVLMs, we find that LVLMs suffer from the hallucination problem, i.e. they tend to generate objects that are inconsistent with the target images in the descriptions. To investigate it, this work presents the first systematic study on object hallucination of LVLMs. We conduct the evaluation experiments on several representative LVLMs, an"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"LVLMs mostly suffer from severe object hallucination issue... objects that frequently occur in the visual instructions or co-occur with the image objects, are obviously prone to be hallucinated by LVLMs... our POPE can evaluate the object hallucination in a more stable and flexible way.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the selected representative LVLMs and visual instruction datasets are sufficiently typical of the broader class of models, and that the polling queries in POPE do not introduce new systematic biases in measuring hallucination.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Large vision-language models exhibit severe object hallucination that varies with training instructions, and the proposed POPE polling method evaluates it more stably and flexibly than prior approaches.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Large vision-language models often describe objects absent from the given image, especially those frequent in instructions or co-occurring with visible items.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"ac61f88496c175c8cae38ec306165295eb469e5b6daaa2714aa7f97bc729e6a4"},"source":{"id":"2305.10355","kind":"arxiv","version":3},"verdict":{"id":"018b7770-e675-418b-841c-c3d57ce286df","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T13:38:09.947713Z","strongest_claim":"LVLMs mostly suffer from severe object hallucination issue... objects that frequently occur in the visual instructions or co-occur with the image objects, are obviously prone to be hallucinated by LVLMs... our POPE can evaluate the object hallucination in a more stable and flexible way.","one_line_summary":"Large vision-language models exhibit severe object hallucination that varies with training instructions, and the proposed POPE polling method evaluates it more stably and flexibly than prior approaches.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the selected representative LVLMs and visual instruction datasets are sufficiently typical of the broader class of models, and that the polling queries in POPE do not introduce new systematic biases in measuring hallucination.","pith_extraction_headline":"Large vision-language models often describe objects absent from the given image, especially those frequent in instructions or co-occurring with visible items."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2305.10355/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":40,"sample":[{"doi":"10.1109/iccv.2019.00904","year":2019,"title":"nocaps: novel object captioning at scale , url=","work_id":"906d0d18-f28e-4142-8d0a-251ea3dde04a","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2022,"title":"Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \\' e n Simonyan","work_id":"24b0d6c0-c337-4551-8013-ce640e0b131a","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2015,"title":"Lawrence Zitnick, and Devi Parikh","work_id":"27ece143-2394-41c7-a24f-3e240c2ff2f7","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.1109/iccv.2015.279","year":2015,"title":"Lawrence Zitnick, and Devi Parikh","work_id":"acf6495e-ae4c-4644-a52d-ff5e1c2ca351","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2023,"title":"Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond","work_id":"cbc2bb21-b6bb-46c0-80bf-107e195ffe10","ref_index":5,"cited_arxiv_id":"2308.12966","is_internal_anchor":true}],"resolved_work":40,"snapshot_sha256":"514a31a0223f5ffb9dfc20f2c07605172ca17f90060be62faed5bb9c0be10516","internal_anchors":12},"formal_canon":{"evidence_count":2,"snapshot_sha256":"abc18560f227d042c33c1ed0bbc904d496b8d8e0da355126fbb20482aeb16438"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2305.10355","created_at":"2026-07-05T07:05:08.636640+00:00"},{"alias_kind":"arxiv_version","alias_value":"2305.10355v3","created_at":"2026-07-05T07:05:08.636640+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2305.10355","created_at":"2026-07-05T07:05:08.636640+00:00"},{"alias_kind":"pith_short_12","alias_value":"555TZ7WGIQES","created_at":"2026-07-05T07:05:08.636640+00:00"},{"alias_kind":"pith_short_16","alias_value":"555TZ7WGIQESUA54","created_at":"2026-07-05T07:05:08.636640+00:00"},{"alias_kind":"pith_short_8","alias_value":"555TZ7WG","created_at":"2026-07-05T07:05:08.636640+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":118,"internal_anchor_count":118,"sample":[{"citing_arxiv_id":"2607.08194","citing_title":"Dive Into the Implicit Biases of Low-rank Vision-language Alignment","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26529","citing_title":"The inattentional gap in task conditioned AI models that omit otherwise reportable safety critical signals","ref_index":32,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26587","citing_title":"SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22352","citing_title":"On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning","ref_index":21,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21734","citing_title":"HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning","ref_index":273,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17030","citing_title":"Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation","ref_index":54,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17118","citing_title":"MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs","ref_index":7,"is_internal_anchor":true},{"citing_arxiv_id":"2606.13289","citing_title":"HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers","ref_index":91,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11096","citing_title":"IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11188","citing_title":"ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations","ref_index":44,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07861","citing_title":"The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models","ref_index":30,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07647","citing_title":"Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03569","citing_title":"When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03713","citing_title":"Investigating Adversarial Robustness of Multi-modal Large Language Models","ref_index":28,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03376","citing_title":"P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization","ref_index":89,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07639","citing_title":"MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention","ref_index":31,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27596","citing_title":"Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding","ref_index":83,"is_internal_anchor":true},{"citing_arxiv_id":"2605.18160","citing_title":"Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2605.24602","citing_title":"Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.29812","citing_title":"Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2605.25343","citing_title":"Toward Native Multimodal Modeling: A Roadmap","ref_index":271,"is_internal_anchor":true},{"citing_arxiv_id":"2605.26601","citing_title":"FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2605.26761","citing_title":"Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning","ref_index":30,"is_internal_anchor":true},{"citing_arxiv_id":"2605.30912","citing_title":"Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2306.00978","citing_title":"AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration","ref_index":22,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4","json":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4.json","graph_json":"https://pith.science/api/pith-number/555TZ7WGIQESUA54MQKZTEW5X4/graph.json","events_json":"https://pith.science/api/pith-number/555TZ7WGIQESUA54MQKZTEW5X4/events.json","paper":"https://pith.science/paper/555TZ7WG"},"agent_actions":{"view_html":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4","download_json":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4.json","view_paper":"https://pith.science/paper/555TZ7WG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2305.10355&json=true","fetch_graph":"https://pith.science/api/pith-number/555TZ7WGIQESUA54MQKZTEW5X4/graph.json","fetch_events":"https://pith.science/api/pith-number/555TZ7WGIQESUA54MQKZTEW5X4/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4/action/timestamp_anchor","attest_storage":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4/action/storage_attestation","attest_author":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4/action/author_attestation","sign_citation":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4/action/citation_signature","submit_replication":"https://pith.science/pith/555TZ7WGIQESUA54MQKZTEW5X4/action/replication_record"}},"created_at":"2026-07-05T07:05:08.636640+00:00","updated_at":"2026-07-05T07:05:08.636640+00:00"}