{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:FSOSMCLVFPN6LZ5EB5GWMS5CTW","short_pith_number":"pith:FSOSMCLV","schema_version":"1.0","canonical_sha256":"2c9d2609752bdbe5e7a40f4d664ba29d9ea3ba03f784127e94e72f57a03fc4d3","source":{"kind":"arxiv","id":"2406.10900","version":2},"attestation_state":"computed","paper":{"title":"AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Abhinav Shrivastava, Dianqi Li, Dinesh Manocha, Furong Huang, Jordan Lee Boyd-Graber, Ruiqi Xian, Shuaiyi Huang, Tianrui Guan, Tianyi Zhou, Xiaoyu Liu, Xijun Wang, Xiyang Wu","submitted_at":"2024-06-16T11:44:43Z","abstract_excerpt":"Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucinations, they often rely on hand-crafted corner cases whose failure patterns may not generalize well. Additionally, fine-tuning on these examples could undermine their validity. To address this, we aim to scale up the number of cases through an automated approach, reducing human bias in crafting such "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2406.10900","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2024-06-16T11:44:43Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"546d96d696c4e1d590b84d098341ae3648321f1436c4176b01e7d835eb3eed98","abstract_canon_sha256":"325753ce914103d11a6a0284e69c424336e325a9b0d4dabc4715ed1e5d431f69"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:17:43.449942Z","signature_b64":"x8yedP0bb93+lCyTQQM2q6v/kUqbVPoG1GGDwhMOBDwoQx4Nb6qWeji8dXantItVL//HOhOFFl+ppFJ90gihDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2c9d2609752bdbe5e7a40f4d664ba29d9ea3ba03f784127e94e72f57a03fc4d3","last_reissued_at":"2026-07-05T09:17:43.449463Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:17:43.449463Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Abhinav Shrivastava, Dianqi Li, Dinesh Manocha, Furong Huang, Jordan Lee Boyd-Graber, Ruiqi Xian, Shuaiyi Huang, Tianrui Guan, Tianyi Zhou, Xiaoyu Liu, Xijun Wang, Xiyang Wu","submitted_at":"2024-06-16T11:44:43Z","abstract_excerpt":"Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some benchmarks have been developed to investigate LVLM hallucinations, they often rely on hand-crafted corner cases whose failure patterns may not generalize well. Additionally, fine-tuning on these examples could undermine their validity. To address this, we aim to scale up the number of cases through an automated approach, reducing human bias in crafting such "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2406.10900","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2406.10900/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2406.10900","created_at":"2026-07-05T09:17:43.449516+00:00"},{"alias_kind":"arxiv_version","alias_value":"2406.10900v2","created_at":"2026-07-05T09:17:43.449516+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2406.10900","created_at":"2026-07-05T09:17:43.449516+00:00"},{"alias_kind":"pith_short_12","alias_value":"FSOSMCLVFPN6","created_at":"2026-07-05T09:17:43.449516+00:00"},{"alias_kind":"pith_short_16","alias_value":"FSOSMCLVFPN6LZ5E","created_at":"2026-07-05T09:17:43.449516+00:00"},{"alias_kind":"pith_short_8","alias_value":"FSOSMCLV","created_at":"2026-07-05T09:17:43.449516+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":3,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.17953","citing_title":"MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2606.31933","citing_title":"No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs","ref_index":77,"is_internal_anchor":false},{"citing_arxiv_id":"2511.18373","citing_title":"MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models","ref_index":50,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW","json":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW.json","graph_json":"https://pith.science/api/pith-number/FSOSMCLVFPN6LZ5EB5GWMS5CTW/graph.json","events_json":"https://pith.science/api/pith-number/FSOSMCLVFPN6LZ5EB5GWMS5CTW/events.json","paper":"https://pith.science/paper/FSOSMCLV"},"agent_actions":{"view_html":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW","download_json":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW.json","view_paper":"https://pith.science/paper/FSOSMCLV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2406.10900&json=true","fetch_graph":"https://pith.science/api/pith-number/FSOSMCLVFPN6LZ5EB5GWMS5CTW/graph.json","fetch_events":"https://pith.science/api/pith-number/FSOSMCLVFPN6LZ5EB5GWMS5CTW/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW/action/timestamp_anchor","attest_storage":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW/action/storage_attestation","attest_author":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW/action/author_attestation","sign_citation":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW/action/citation_signature","submit_replication":"https://pith.science/pith/FSOSMCLVFPN6LZ5EB5GWMS5CTW/action/replication_record"}},"created_at":"2026-07-05T09:17:43.449516+00:00","updated_at":"2026-07-05T09:17:43.449516+00:00"}