{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:INN6OVCHQXHR3LRXLREGN2XZ3B","short_pith_number":"pith:INN6OVCH","schema_version":"1.0","canonical_sha256":"435be7544785cf1dae375c4866eaf9d868d612d2a2eb1d3b9b894029c1a9d270","source":{"kind":"arxiv","id":"2410.03577","version":2},"attestation_state":"computed","paper":{"title":"Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Chang Tang, Jia Liu, Junkai Chen, Kening Zheng, Peijie Jiang, Sirui Huang, Xin Zou, Xuming Hu, Yibo Yan, Yizhou Wang, Yuanhuiyi Lyu","submitted_at":"2024-10-04T16:30:54Z","abstract_excerpt":"Despite their impressive capabilities, multimodal large language models (MLLMs) are prone to hallucinations, i.e., the generated content that is nonsensical or unfaithful to input sources. Unlike in LLMs, hallucinations in MLLMs often stem from the sensitivity of text decoder to visual tokens, leading to a phenomenon akin to \"amnesia\" about visual information. To address this issue, we propose MemVR, a novel decoding paradigm inspired by common cognition: when the memory of an image seen the moment before is forgotten, people will look at it again for factual answers. Following this principle,"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.03577","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2024-10-04T16:30:54Z","cross_cats_sorted":[],"title_canon_sha256":"492bbc0d12721b0418afef0819e13ddd3ec084f557330ef05c69e7a413bb05ce","abstract_canon_sha256":"d8493edf96e248226182bf91e767b392491d10529fcdc8fbd538cc8352262d7c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:00:00.738630Z","signature_b64":"Ov6OG4IcFkJtBwRt1WMalvW0AXqT8uNrfesVzq/2I0Vr7zX6RwyfYOF7viGiYA5hvVZPyUiUrK/0eMaQdLjuBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"435be7544785cf1dae375c4866eaf9d868d612d2a2eb1d3b9b894029c1a9d270","last_reissued_at":"2026-07-05T11:00:00.738124Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:00:00.738124Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Chang Tang, Jia Liu, Junkai Chen, Kening Zheng, Peijie Jiang, Sirui Huang, Xin Zou, Xuming Hu, Yibo Yan, Yizhou Wang, Yuanhuiyi Lyu","submitted_at":"2024-10-04T16:30:54Z","abstract_excerpt":"Despite their impressive capabilities, multimodal large language models (MLLMs) are prone to hallucinations, i.e., the generated content that is nonsensical or unfaithful to input sources. Unlike in LLMs, hallucinations in MLLMs often stem from the sensitivity of text decoder to visual tokens, leading to a phenomenon akin to \"amnesia\" about visual information. To address this issue, we propose MemVR, a novel decoding paradigm inspired by common cognition: when the memory of an image seen the moment before is forgotten, people will look at it again for factual answers. Following this principle,"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.03577","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.03577/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.03577","created_at":"2026-07-05T11:00:00.738181+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.03577v2","created_at":"2026-07-05T11:00:00.738181+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.03577","created_at":"2026-07-05T11:00:00.738181+00:00"},{"alias_kind":"pith_short_12","alias_value":"INN6OVCHQXHR","created_at":"2026-07-05T11:00:00.738181+00:00"},{"alias_kind":"pith_short_16","alias_value":"INN6OVCHQXHR3LRX","created_at":"2026-07-05T11:00:00.738181+00:00"},{"alias_kind":"pith_short_8","alias_value":"INN6OVCH","created_at":"2026-07-05T11:00:00.738181+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":18,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.13441","citing_title":"Why Sampling Is Not Choosing: Intentionality, Agency, and Moral Responsibility in Large Language Models","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11868","citing_title":"MemNovo: Look Back at the Spectrum for Balanced De Novo Peptide Sequencing from Mass Spectrometry","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2606.27596","citing_title":"Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2606.31054","citing_title":"ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29812","citing_title":"Consistency as Inductive Bias: Learning Cross-View Invariance for Robust Multimodal Reasoning","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00105","citing_title":"Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2502.02871","citing_title":"Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning","ref_index":269,"is_internal_anchor":false},{"citing_arxiv_id":"2511.19972","citing_title":"Boosting Reasoning in Large Multimodal Models via Activation Replay","ref_index":69,"is_internal_anchor":false},{"citing_arxiv_id":"2512.12623","citing_title":"Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2604.03179","citing_title":"Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2604.03045","citing_title":"STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10622","citing_title":"Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination","ref_index":54,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08145","citing_title":"Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00814","citing_title":"Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs","ref_index":98,"is_internal_anchor":false},{"citing_arxiv_id":"2604.12582","citing_title":"Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01766","citing_title":"Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00814","citing_title":"Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs","ref_index":98,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21027","citing_title":"HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering","ref_index":292,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B","json":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B.json","graph_json":"https://pith.science/api/pith-number/INN6OVCHQXHR3LRXLREGN2XZ3B/graph.json","events_json":"https://pith.science/api/pith-number/INN6OVCHQXHR3LRXLREGN2XZ3B/events.json","paper":"https://pith.science/paper/INN6OVCH"},"agent_actions":{"view_html":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B","download_json":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B.json","view_paper":"https://pith.science/paper/INN6OVCH","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.03577&json=true","fetch_graph":"https://pith.science/api/pith-number/INN6OVCHQXHR3LRXLREGN2XZ3B/graph.json","fetch_events":"https://pith.science/api/pith-number/INN6OVCHQXHR3LRXLREGN2XZ3B/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B/action/timestamp_anchor","attest_storage":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B/action/storage_attestation","attest_author":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B/action/author_attestation","sign_citation":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B/action/citation_signature","submit_replication":"https://pith.science/pith/INN6OVCHQXHR3LRXLREGN2XZ3B/action/replication_record"}},"created_at":"2026-07-05T11:00:00.738181+00:00","updated_at":"2026-07-05T11:00:00.738181+00:00"}