{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:2HW5FCXTIWQH5NWTGLLP7Q3EXO","short_pith_number":"pith:2HW5FCXT","schema_version":"1.0","canonical_sha256":"d1edd28af345a07eb6d332d6ffc364bba0604da0148be74743574683db7e0822","source":{"kind":"arxiv","id":"2310.03025","version":2},"attestation_state":"computed","paper":{"title":"Retrieval meets Long Context Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.IR","cs.LG"],"primary_cat":"cs.CL","authors_text":"Bryan Catanzaro, Chen Zhu, Evelina Bakhturina, Lawrence McAfee, Mohammad Shoeybi, Peng Xu, Sandeep Subramanian, Wei Ping, Xianchao Wu, Zihan Liu","submitted_at":"2023-10-04T17:59:41Z","abstract_excerpt":"Extending the context window of large language models (LLMs) is getting popular recently, while the solution of augmenting LLMs with retrieval has existed for years. The natural questions are: i) Retrieval-augmentation versus long context window, which one is better for downstream tasks? ii) Can both methods be combined to get the best of both worlds? In this work, we answer these questions by studying both solutions using two state-of-the-art pretrained LLMs, i.e., a proprietary 43B GPT and Llama2-70B. Perhaps surprisingly, we find that LLM with 4K context window using simple retrieval-augmen"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.03025","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2023-10-04T17:59:41Z","cross_cats_sorted":["cs.AI","cs.IR","cs.LG"],"title_canon_sha256":"973f7cad035fc1ff755da322f09784c3a49dab8b9d6c799d5a4b89b3fca3080a","abstract_canon_sha256":"fedc53f7b8cdacc50fbb55c1944294779d05bbda1e344527b138ef040140cc79"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:36:24.222382Z","signature_b64":"j+CvLn+THFUYQhOa8BVzykp945f+rfR5NRTcF4y6iIMsL0sqOgXWcxUPNaGfCr12Qn7MFYo5lBKPhVEsCnltCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"d1edd28af345a07eb6d332d6ffc364bba0604da0148be74743574683db7e0822","last_reissued_at":"2026-07-05T07:36:24.221933Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:36:24.221933Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Retrieval meets Long Context Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.IR","cs.LG"],"primary_cat":"cs.CL","authors_text":"Bryan Catanzaro, Chen Zhu, Evelina Bakhturina, Lawrence McAfee, Mohammad Shoeybi, Peng Xu, Sandeep Subramanian, Wei Ping, Xianchao Wu, Zihan Liu","submitted_at":"2023-10-04T17:59:41Z","abstract_excerpt":"Extending the context window of large language models (LLMs) is getting popular recently, while the solution of augmenting LLMs with retrieval has existed for years. The natural questions are: i) Retrieval-augmentation versus long context window, which one is better for downstream tasks? ii) Can both methods be combined to get the best of both worlds? In this work, we answer these questions by studying both solutions using two state-of-the-art pretrained LLMs, i.e., a proprietary 43B GPT and Llama2-70B. Perhaps surprisingly, we find that LLM with 4K context window using simple retrieval-augmen"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.03025","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.03025/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.03025","created_at":"2026-07-05T07:36:24.221989+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.03025v2","created_at":"2026-07-05T07:36:24.221989+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.03025","created_at":"2026-07-05T07:36:24.221989+00:00"},{"alias_kind":"pith_short_12","alias_value":"2HW5FCXTIWQH","created_at":"2026-07-05T07:36:24.221989+00:00"},{"alias_kind":"pith_short_16","alias_value":"2HW5FCXTIWQH5NWT","created_at":"2026-07-05T07:36:24.221989+00:00"},{"alias_kind":"pith_short_8","alias_value":"2HW5FCXT","created_at":"2026-07-05T07:36:24.221989+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2605.01342","citing_title":"Don't Stir the Pot! Authorized Vector Data Retrieval via Access-Aware Indexing","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26100","citing_title":"Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2312.10997","citing_title":"Retrieval-Augmented Generation for Large Language Models: A Survey","ref_index":170,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01342","citing_title":"Don't Stir the Pot! Authorized Vector Data Retrieval via Access-Aware Indexing","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06330","citing_title":"Fine-Tuning Small Language Models for Solution-Oriented Windows Event Log Analysis","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01342","citing_title":"Don't Stir the Pot! Authorized Vector Data Retrieval via Access-Aware Indexing","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2604.07041","citing_title":"AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views","ref_index":48,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO","json":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO.json","graph_json":"https://pith.science/api/pith-number/2HW5FCXTIWQH5NWTGLLP7Q3EXO/graph.json","events_json":"https://pith.science/api/pith-number/2HW5FCXTIWQH5NWTGLLP7Q3EXO/events.json","paper":"https://pith.science/paper/2HW5FCXT"},"agent_actions":{"view_html":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO","download_json":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO.json","view_paper":"https://pith.science/paper/2HW5FCXT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.03025&json=true","fetch_graph":"https://pith.science/api/pith-number/2HW5FCXTIWQH5NWTGLLP7Q3EXO/graph.json","fetch_events":"https://pith.science/api/pith-number/2HW5FCXTIWQH5NWTGLLP7Q3EXO/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO/action/timestamp_anchor","attest_storage":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO/action/storage_attestation","attest_author":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO/action/author_attestation","sign_citation":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO/action/citation_signature","submit_replication":"https://pith.science/pith/2HW5FCXTIWQH5NWTGLLP7Q3EXO/action/replication_record"}},"created_at":"2026-07-05T07:36:24.221989+00:00","updated_at":"2026-07-05T07:36:24.221989+00:00"}