{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:7ZMX4MQX43AU6MXCRS273HP72A","short_pith_number":"pith:7ZMX4MQX","schema_version":"1.0","canonical_sha256":"fe597e3217e6c14f32e28cb5fd9dffd02eea2bd22b2806d04aa219574220b1ef","source":{"kind":"arxiv","id":"2211.12561","version":2},"attestation_state":"computed","paper":{"title":"Retrieval-Augmented Multimodal Language Modeling","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL","cs.LG"],"primary_cat":"cs.CV","authors_text":"Armen Aghajanyan, Jure Leskovec, Luke Zettlemoyer, Michihiro Yasunaga, Mike Lewis, Percy Liang, Rich James, Weijia Shi, Wen-tau Yih","submitted_at":"2022-11-22T20:26:44Z","abstract_excerpt":"Recent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all learned knowledge (e.g., the appearance of the Eiffel Tower) in the model parameters, requiring increasingly larger models and training data to capture more knowledge. To integrate knowledge in a more scalable and modular way, we propose a retrieval-augmented multimodal model, which enables a base multimodal model (generator) to refer to relevant text and images fetched by a retriever from external memory (e.g., documents on the web). "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2211.12561","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2022-11-22T20:26:44Z","cross_cats_sorted":["cs.CL","cs.LG"],"title_canon_sha256":"2db65baa3f5058b9bcb69cfb0cc36148a4a9c9417be9e02f9f798828a7fc0bd4","abstract_canon_sha256":"a45f8bd32ebc4991e7e063dd4ba7d0411721356c43185d291921cfe62d927af3"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:17:42.690885Z","signature_b64":"vzb9Jv19fves1tQ4zDtWhfv2JGfP2/OvmVvsdSsNrtsB7OzxyTScHbARusJ2qp14oV6uqg0YnYwDaeIAgJsyAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"fe597e3217e6c14f32e28cb5fd9dffd02eea2bd22b2806d04aa219574220b1ef","last_reissued_at":"2026-07-05T06:17:42.690444Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:17:42.690444Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Retrieval-Augmented Multimodal Language Modeling","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL","cs.LG"],"primary_cat":"cs.CV","authors_text":"Armen Aghajanyan, Jure Leskovec, Luke Zettlemoyer, Michihiro Yasunaga, Mike Lewis, Percy Liang, Rich James, Weijia Shi, Wen-tau Yih","submitted_at":"2022-11-22T20:26:44Z","abstract_excerpt":"Recent multimodal models such as DALL-E and CM3 have achieved remarkable progress in text-to-image and image-to-text generation. However, these models store all learned knowledge (e.g., the appearance of the Eiffel Tower) in the model parameters, requiring increasingly larger models and training data to capture more knowledge. To integrate knowledge in a more scalable and modular way, we propose a retrieval-augmented multimodal model, which enables a base multimodal model (generator) to refer to relevant text and images fetched by a retriever from external memory (e.g., documents on the web). "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2211.12561","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2211.12561/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2211.12561","created_at":"2026-07-05T06:17:42.690500+00:00"},{"alias_kind":"arxiv_version","alias_value":"2211.12561v2","created_at":"2026-07-05T06:17:42.690500+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2211.12561","created_at":"2026-07-05T06:17:42.690500+00:00"},{"alias_kind":"pith_short_12","alias_value":"7ZMX4MQX43AU","created_at":"2026-07-05T06:17:42.690500+00:00"},{"alias_kind":"pith_short_16","alias_value":"7ZMX4MQX43AU6MXC","created_at":"2026-07-05T06:17:42.690500+00:00"},{"alias_kind":"pith_short_8","alias_value":"7ZMX4MQX","created_at":"2026-07-05T06:17:42.690500+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":14,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25343","citing_title":"Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity","ref_index":75,"is_internal_anchor":false},{"citing_arxiv_id":"2606.26916","citing_title":"PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation","ref_index":81,"is_internal_anchor":false},{"citing_arxiv_id":"2606.25343","citing_title":"Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity","ref_index":75,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20173","citing_title":"Qiskit Code Migration with LLMs","ref_index":100,"is_internal_anchor":false},{"citing_arxiv_id":"2312.10997","citing_title":"Retrieval-Augmented Generation for Large Language Models: A Survey","ref_index":176,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17118","citing_title":"Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2408.04840","citing_title":"mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2506.04565","citing_title":"From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems","ref_index":212,"is_internal_anchor":false},{"citing_arxiv_id":"2511.13415","citing_title":"Attention Grounded Enhancement for Visual Document Retrieval","ref_index":59,"is_internal_anchor":false},{"citing_arxiv_id":"2301.12652","citing_title":"REPLUG: Retrieval-Augmented Black-Box Language Models","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2302.14045","citing_title":"Language Is Not All You Need: Aligning Perception with Language Models","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2310.11511","citing_title":"Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection","ref_index":106,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00814","citing_title":"Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs","ref_index":87,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00814","citing_title":"Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs","ref_index":87,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A","json":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A.json","graph_json":"https://pith.science/api/pith-number/7ZMX4MQX43AU6MXCRS273HP72A/graph.json","events_json":"https://pith.science/api/pith-number/7ZMX4MQX43AU6MXCRS273HP72A/events.json","paper":"https://pith.science/paper/7ZMX4MQX"},"agent_actions":{"view_html":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A","download_json":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A.json","view_paper":"https://pith.science/paper/7ZMX4MQX","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2211.12561&json=true","fetch_graph":"https://pith.science/api/pith-number/7ZMX4MQX43AU6MXCRS273HP72A/graph.json","fetch_events":"https://pith.science/api/pith-number/7ZMX4MQX43AU6MXCRS273HP72A/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A/action/timestamp_anchor","attest_storage":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A/action/storage_attestation","attest_author":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A/action/author_attestation","sign_citation":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A/action/citation_signature","submit_replication":"https://pith.science/pith/7ZMX4MQX43AU6MXCRS273HP72A/action/replication_record"}},"created_at":"2026-07-05T06:17:42.690500+00:00","updated_at":"2026-07-05T06:17:42.690500+00:00"}