{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:TQVV4L6WXXMLVLVVEIXC42SYJW","short_pith_number":"pith:TQVV4L6W","schema_version":"1.0","canonical_sha256":"9c2b5e2fd6bdd8baaeb5222e2e6a584d8aadd6f61e8a10ca9884f48622488c87","source":{"kind":"arxiv","id":"2411.09968","version":1},"attestation_state":"computed","paper":{"title":"Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Chaochen Gu, Chen Shen, Hao Cheng, Jieping Ye, Kaijie Wu, Shaotian Yan, Xiaofeng Zhang, Xiaosong Yuan, Yihao Quan","submitted_at":"2024-11-15T05:51:29Z","abstract_excerpt":"The hallucination problem in multimodal large language models (MLLMs) remains a common issue. Although image tokens occupy a majority of the input sequence of MLLMs, there is limited research to explore the relationship between image tokens and hallucinations. In this paper, we analyze the distribution of attention scores for image tokens across each layer and head of the model, revealing an intriguing and common phenomenon: most hallucinations are closely linked to the pattern of attention sinks in the self-attention matrix of image tokens, where shallow layers exhibit dense attention sinks a"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2411.09968","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-sa/4.0/","primary_cat":"cs.CV","submitted_at":"2024-11-15T05:51:29Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"a2ff753811b80b57d2c3b06bd40eebaecb02f6461c7843d041565a4b51ecc98f","abstract_canon_sha256":"ad8256e31126c97f77b4cd8265a6a9e9d0a7e242323df3eb69d5e715c46a6b35"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:35:56.516878Z","signature_b64":"TUQFl05vJavHNvAuRpSUgGMtM2B9iiLpduuw3WhHSv4LcSCws4m9EDyMNv4VnxONV4XsjgIAsNYX2K7QvQasBA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"9c2b5e2fd6bdd8baaeb5222e2e6a584d8aadd6f61e8a10ca9884f48622488c87","last_reissued_at":"2026-07-05T09:35:56.516381Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:35:56.516381Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Chaochen Gu, Chen Shen, Hao Cheng, Jieping Ye, Kaijie Wu, Shaotian Yan, Xiaofeng Zhang, Xiaosong Yuan, Yihao Quan","submitted_at":"2024-11-15T05:51:29Z","abstract_excerpt":"The hallucination problem in multimodal large language models (MLLMs) remains a common issue. Although image tokens occupy a majority of the input sequence of MLLMs, there is limited research to explore the relationship between image tokens and hallucinations. In this paper, we analyze the distribution of attention scores for image tokens across each layer and head of the model, revealing an intriguing and common phenomenon: most hallucinations are closely linked to the pattern of attention sinks in the self-attention matrix of image tokens, where shallow layers exhibit dense attention sinks a"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2411.09968","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2411.09968/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2411.09968","created_at":"2026-07-05T09:35:56.516445+00:00"},{"alias_kind":"arxiv_version","alias_value":"2411.09968v1","created_at":"2026-07-05T09:35:56.516445+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2411.09968","created_at":"2026-07-05T09:35:56.516445+00:00"},{"alias_kind":"pith_short_12","alias_value":"TQVV4L6WXXML","created_at":"2026-07-05T09:35:56.516445+00:00"},{"alias_kind":"pith_short_16","alias_value":"TQVV4L6WXXMLVLVV","created_at":"2026-07-05T09:35:56.516445+00:00"},{"alias_kind":"pith_short_8","alias_value":"TQVV4L6W","created_at":"2026-07-05T09:35:56.516445+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":11,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25325","citing_title":"Omni-Perception Policy Optimization for Multimodal Emotion Reasoning","ref_index":76,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10651","citing_title":"Kwai Keye-VL-2.0 Technical Report","ref_index":70,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29805","citing_title":"Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation","ref_index":81,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29805","citing_title":"Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation","ref_index":81,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25799","citing_title":"Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2505.21472","citing_title":"Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11559","citing_title":"When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04874","citing_title":"Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2404.18930","citing_title":"Hallucination of Multimodal Large Language Models: A Survey","ref_index":213,"is_internal_anchor":false},{"citing_arxiv_id":"2604.10098","citing_title":"Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation","ref_index":115,"is_internal_anchor":false},{"citing_arxiv_id":"2604.06393","citing_title":"ART: Attention Replacement Technique to Improve Factuality in LLMs","ref_index":15,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW","json":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW.json","graph_json":"https://pith.science/api/pith-number/TQVV4L6WXXMLVLVVEIXC42SYJW/graph.json","events_json":"https://pith.science/api/pith-number/TQVV4L6WXXMLVLVVEIXC42SYJW/events.json","paper":"https://pith.science/paper/TQVV4L6W"},"agent_actions":{"view_html":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW","download_json":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW.json","view_paper":"https://pith.science/paper/TQVV4L6W","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2411.09968&json=true","fetch_graph":"https://pith.science/api/pith-number/TQVV4L6WXXMLVLVVEIXC42SYJW/graph.json","fetch_events":"https://pith.science/api/pith-number/TQVV4L6WXXMLVLVVEIXC42SYJW/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW/action/timestamp_anchor","attest_storage":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW/action/storage_attestation","attest_author":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW/action/author_attestation","sign_citation":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW/action/citation_signature","submit_replication":"https://pith.science/pith/TQVV4L6WXXMLVLVVEIXC42SYJW/action/replication_record"}},"created_at":"2026-07-05T09:35:56.516445+00:00","updated_at":"2026-07-05T09:35:56.516445+00:00"}