{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:RBVDWFGDGXUT3JXB3JHSDBWWDK","short_pith_number":"pith:RBVDWFGD","schema_version":"1.0","canonical_sha256":"886a3b14c335e93da6e1da4f2186d61aaf20b26d351a6273f036a0203c6ccc95","source":{"kind":"arxiv","id":"2411.12644","version":3},"attestation_state":"computed","paper":{"title":"CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.SE","authors_text":"Caiming Xiong, Rui Meng, Semih Yavuz, Shafiq Joty, Silvio Savarese, Ye Liu, Yingbo Zhou","submitted_at":"2024-11-19T16:54:45Z","abstract_excerpt":"Despite the success of text retrieval in many NLP tasks, code retrieval remains a largely underexplored area. Most text retrieval systems are tailored for natural language queries, often neglecting the specific challenges of retrieving code. This gap leaves existing models unable to effectively capture the diversity of programming languages and tasks across different domains, highlighting the need for more focused research in code retrieval. To address this, we introduce CodeXEmbed, a family of large-scale code embedding models ranging from 400M to 7B parameters. Our novel training pipeline un"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2411.12644","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.SE","submitted_at":"2024-11-19T16:54:45Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"4e61fdb1bf92d837ea0bede9e72ab0595b5b507673bf7baafa81d92b6c9b02c3","abstract_canon_sha256":"855e560d54f414238159ee2cddb1f63ffea56fe3c2e2e1c2cc424b885236992e"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:50:29.055999Z","signature_b64":"eyYQvzrr8NTEAJgbapcwmpdh51wVT33KjKhDdhJML20SQpqqJKI59pEO7HezJWJ5TtuFHsCrjzXR5MPMZBDTAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"886a3b14c335e93da6e1da4f2186d61aaf20b26d351a6273f036a0203c6ccc95","last_reissued_at":"2026-07-05T11:50:29.055524Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:50:29.055524Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.SE","authors_text":"Caiming Xiong, Rui Meng, Semih Yavuz, Shafiq Joty, Silvio Savarese, Ye Liu, Yingbo Zhou","submitted_at":"2024-11-19T16:54:45Z","abstract_excerpt":"Despite the success of text retrieval in many NLP tasks, code retrieval remains a largely underexplored area. Most text retrieval systems are tailored for natural language queries, often neglecting the specific challenges of retrieving code. This gap leaves existing models unable to effectively capture the diversity of programming languages and tasks across different domains, highlighting the need for more focused research in code retrieval. To address this, we introduce CodeXEmbed, a family of large-scale code embedding models ranging from 400M to 7B parameters. Our novel training pipeline un"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2411.12644","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2411.12644/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2411.12644","created_at":"2026-07-05T11:50:29.055589+00:00"},{"alias_kind":"arxiv_version","alias_value":"2411.12644v3","created_at":"2026-07-05T11:50:29.055589+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2411.12644","created_at":"2026-07-05T11:50:29.055589+00:00"},{"alias_kind":"pith_short_12","alias_value":"RBVDWFGDGXUT","created_at":"2026-07-05T11:50:29.055589+00:00"},{"alias_kind":"pith_short_16","alias_value":"RBVDWFGDGXUT3JXB","created_at":"2026-07-05T11:50:29.055589+00:00"},{"alias_kind":"pith_short_8","alias_value":"RBVDWFGD","created_at":"2026-07-05T11:50:29.055589+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.31725","citing_title":"Do Machines Struggle Where Humans Do? LLM and Human Comprehension of Obfuscated Code","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.27787","citing_title":"Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering Systems","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14503","citing_title":"Not All RAGs Are Created Equal: A Component-Wise Empirical Study for Software Engineering Tasks","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08299","citing_title":"Do not copy and paste! Rewriting strategies for code retrieval","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08083","citing_title":"Can LLMs Deobfuscate Binary Code? A Systematic Analysis of Large Language Models into Pseudocode Deobfuscation","ref_index":75,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK","json":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK.json","graph_json":"https://pith.science/api/pith-number/RBVDWFGDGXUT3JXB3JHSDBWWDK/graph.json","events_json":"https://pith.science/api/pith-number/RBVDWFGDGXUT3JXB3JHSDBWWDK/events.json","paper":"https://pith.science/paper/RBVDWFGD"},"agent_actions":{"view_html":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK","download_json":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK.json","view_paper":"https://pith.science/paper/RBVDWFGD","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2411.12644&json=true","fetch_graph":"https://pith.science/api/pith-number/RBVDWFGDGXUT3JXB3JHSDBWWDK/graph.json","fetch_events":"https://pith.science/api/pith-number/RBVDWFGDGXUT3JXB3JHSDBWWDK/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK/action/timestamp_anchor","attest_storage":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK/action/storage_attestation","attest_author":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK/action/author_attestation","sign_citation":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK/action/citation_signature","submit_replication":"https://pith.science/pith/RBVDWFGDGXUT3JXB3JHSDBWWDK/action/replication_record"}},"created_at":"2026-07-05T11:50:29.055589+00:00","updated_at":"2026-07-05T11:50:29.055589+00:00"}