{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:JWORBE7CXDRNFLZ72FQFEXSCZW","short_pith_number":"pith:JWORBE7C","schema_version":"1.0","canonical_sha256":"4d9d1093e2b8e2d2af3fd160525e42cdb9ffe602714669a68d1821789080116d","source":{"kind":"arxiv","id":"2402.04248","version":2},"attestation_state":"computed","paper":{"title":"Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Dimitris Papailiopoulos, Jaeseung Park, Jaewoong Cho, Jongho Park, Kangwook Lee, Nayoung Lee, Samet Oymak, Zheyang Xiong","submitted_at":"2024-02-06T18:56:35Z","abstract_excerpt":"State-space models (SSMs), such as Mamba (Gu & Dao, 2023), have been proposed as alternatives to Transformer networks in language modeling, by incorporating gating, convolutions, and input-dependent token selection to mitigate the quadratic cost of multi-head attention. Although SSMs exhibit competitive performance, their in-context learning (ICL) capabilities, a remarkable emergent property of modern language models that enables task execution without parameter optimization, remain underexplored compared to Transformers. In this study, we evaluate the ICL performance of SSMs, focusing on Mamb"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.04248","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-02-06T18:56:35Z","cross_cats_sorted":[],"title_canon_sha256":"ba0771ecf258d1c443a82c59fdea3b2540f51ca37da0ae216114b969152ebf3d","abstract_canon_sha256":"e8bed7bfcfbb597fd30011da299dc9d8942c6d3a56df78b1731296984f23081b"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:11:57.590536Z","signature_b64":"4JiiwjrwJJ68SjiIQ8EasAgqOC+zvD63clCBXInUd3FFDI+E4+L3d+jnS3/bJ28XrAxBwLWpdb2lJNZr+hGUBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4d9d1093e2b8e2d2af3fd160525e42cdb9ffe602714669a68d1821789080116d","last_reissued_at":"2026-07-05T08:11:57.590037Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:11:57.590037Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Dimitris Papailiopoulos, Jaeseung Park, Jaewoong Cho, Jongho Park, Kangwook Lee, Nayoung Lee, Samet Oymak, Zheyang Xiong","submitted_at":"2024-02-06T18:56:35Z","abstract_excerpt":"State-space models (SSMs), such as Mamba (Gu & Dao, 2023), have been proposed as alternatives to Transformer networks in language modeling, by incorporating gating, convolutions, and input-dependent token selection to mitigate the quadratic cost of multi-head attention. Although SSMs exhibit competitive performance, their in-context learning (ICL) capabilities, a remarkable emergent property of modern language models that enables task execution without parameter optimization, remain underexplored compared to Transformers. In this study, we evaluate the ICL performance of SSMs, focusing on Mamb"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.04248","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.04248/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.04248","created_at":"2026-07-05T08:11:57.590097+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.04248v2","created_at":"2026-07-05T08:11:57.590097+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.04248","created_at":"2026-07-05T08:11:57.590097+00:00"},{"alias_kind":"pith_short_12","alias_value":"JWORBE7CXDRN","created_at":"2026-07-05T08:11:57.590097+00:00"},{"alias_kind":"pith_short_16","alias_value":"JWORBE7CXDRNFLZ7","created_at":"2026-07-05T08:11:57.590097+00:00"},{"alias_kind":"pith_short_8","alias_value":"JWORBE7C","created_at":"2026-07-05T08:11:57.590097+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":13,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.24320","citing_title":"ZONOS2 Technical Report","ref_index":90,"is_internal_anchor":false},{"citing_arxiv_id":"2606.26290","citing_title":"SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2606.24320","citing_title":"ZONOS2 Technical Report","ref_index":90,"is_internal_anchor":false},{"citing_arxiv_id":"2511.05963","citing_title":"Next-Latent Prediction Transformers Learn Compact World Models","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2408.01129","citing_title":"A Survey of Mamba","ref_index":142,"is_internal_anchor":false},{"citing_arxiv_id":"2406.07887","citing_title":"An Empirical Study of Mamba-based Language Models","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2602.01651","citing_title":"On the Spatiotemporal Dynamics of Generalization in Neural Networks","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2404.14294","citing_title":"A Survey on Efficient Inference for Large Language Models","ref_index":76,"is_internal_anchor":false},{"citing_arxiv_id":"2603.29069","citing_title":"On the Mirage of Long-Range Dependency, with an Application to Integer Multiplication","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2403.19887","citing_title":"Jamba: A Hybrid Transformer-Mamba Language Model","ref_index":36,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05365","citing_title":"ZAYA1-8B Technical Report","ref_index":64,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01240","citing_title":"Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01240","citing_title":"Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI","ref_index":44,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW","json":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW.json","graph_json":"https://pith.science/api/pith-number/JWORBE7CXDRNFLZ72FQFEXSCZW/graph.json","events_json":"https://pith.science/api/pith-number/JWORBE7CXDRNFLZ72FQFEXSCZW/events.json","paper":"https://pith.science/paper/JWORBE7C"},"agent_actions":{"view_html":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW","download_json":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW.json","view_paper":"https://pith.science/paper/JWORBE7C","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.04248&json=true","fetch_graph":"https://pith.science/api/pith-number/JWORBE7CXDRNFLZ72FQFEXSCZW/graph.json","fetch_events":"https://pith.science/api/pith-number/JWORBE7CXDRNFLZ72FQFEXSCZW/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW/action/timestamp_anchor","attest_storage":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW/action/storage_attestation","attest_author":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW/action/author_attestation","sign_citation":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW/action/citation_signature","submit_replication":"https://pith.science/pith/JWORBE7CXDRNFLZ72FQFEXSCZW/action/replication_record"}},"created_at":"2026-07-05T08:11:57.590097+00:00","updated_at":"2026-07-05T08:11:57.590097+00:00"}