{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:WRERGC4P3ZAJIHNGICYAVYFELE","short_pith_number":"pith:WRERGC4P","schema_version":"1.0","canonical_sha256":"b449130b8fde40941da640b00ae0a45938831f3828c7fe187ab1753a2f8b1cc6","source":{"kind":"arxiv","id":"2407.19584","version":1},"attestation_state":"computed","paper":{"title":"SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Dominic Culver, Etienne Malaboeuf, Gabriel Hautreux, Johanne Charpentier, Malik Boudiaf, Michael Desa, Pierre Colombo, Rui Melo, Sofia Morgado, Telmo Pires","submitted_at":"2024-07-28T20:50:53Z","abstract_excerpt":"In this paper, we introduce SaulLM-54B and SaulLM-141B, two large language models (LLMs) tailored for the legal sector. These models, which feature architectures of 54 billion and 141 billion parameters, respectively, are based on the Mixtral architecture. The development of SaulLM-54B and SaulLM-141B is guided by large-scale domain adaptation, divided into three strategies: (1) the exploitation of continued pretraining involving a base corpus that includes over 540 billion of legal tokens, (2) the implementation of a specialized legal instruction-following protocol, and (3) the alignment of m"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2407.19584","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2024-07-28T20:50:53Z","cross_cats_sorted":[],"title_canon_sha256":"87f14659d38cc0bd842276ebc4104c94426875f0565a31681b9c988a9b02b970","abstract_canon_sha256":"0a26384dc9ae013d230d766e53f5486944d23deab6234606caa147c6b72ccfd6"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:49:41.842826Z","signature_b64":"dyT7mD49Gkt0q3kWImUjDaTUeSegcfw+eSBUIw41RqvMStJP2JT6+vfa/K+71v3RfXjlZtTPtY6uQQQ3/gGRBA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"b449130b8fde40941da640b00ae0a45938831f3828c7fe187ab1753a2f8b1cc6","last_reissued_at":"2026-07-05T08:49:41.842307Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:49:41.842307Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Dominic Culver, Etienne Malaboeuf, Gabriel Hautreux, Johanne Charpentier, Malik Boudiaf, Michael Desa, Pierre Colombo, Rui Melo, Sofia Morgado, Telmo Pires","submitted_at":"2024-07-28T20:50:53Z","abstract_excerpt":"In this paper, we introduce SaulLM-54B and SaulLM-141B, two large language models (LLMs) tailored for the legal sector. These models, which feature architectures of 54 billion and 141 billion parameters, respectively, are based on the Mixtral architecture. The development of SaulLM-54B and SaulLM-141B is guided by large-scale domain adaptation, divided into three strategies: (1) the exploitation of continued pretraining involving a base corpus that includes over 540 billion of legal tokens, (2) the implementation of a specialized legal instruction-following protocol, and (3) the alignment of m"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2407.19584","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2407.19584/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2407.19584","created_at":"2026-07-05T08:49:41.842394+00:00"},{"alias_kind":"arxiv_version","alias_value":"2407.19584v1","created_at":"2026-07-05T08:49:41.842394+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2407.19584","created_at":"2026-07-05T08:49:41.842394+00:00"},{"alias_kind":"pith_short_12","alias_value":"WRERGC4P3ZAJ","created_at":"2026-07-05T08:49:41.842394+00:00"},{"alias_kind":"pith_short_16","alias_value":"WRERGC4P3ZAJIHNG","created_at":"2026-07-05T08:49:41.842394+00:00"},{"alias_kind":"pith_short_8","alias_value":"WRERGC4P","created_at":"2026-07-05T08:49:41.842394+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":2,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.24901","citing_title":"LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24452","citing_title":"Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions","ref_index":7,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE","json":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE.json","graph_json":"https://pith.science/api/pith-number/WRERGC4P3ZAJIHNGICYAVYFELE/graph.json","events_json":"https://pith.science/api/pith-number/WRERGC4P3ZAJIHNGICYAVYFELE/events.json","paper":"https://pith.science/paper/WRERGC4P"},"agent_actions":{"view_html":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE","download_json":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE.json","view_paper":"https://pith.science/paper/WRERGC4P","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2407.19584&json=true","fetch_graph":"https://pith.science/api/pith-number/WRERGC4P3ZAJIHNGICYAVYFELE/graph.json","fetch_events":"https://pith.science/api/pith-number/WRERGC4P3ZAJIHNGICYAVYFELE/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE/action/timestamp_anchor","attest_storage":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE/action/storage_attestation","attest_author":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE/action/author_attestation","sign_citation":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE/action/citation_signature","submit_replication":"https://pith.science/pith/WRERGC4P3ZAJIHNGICYAVYFELE/action/replication_record"}},"created_at":"2026-07-05T08:49:41.842394+00:00","updated_at":"2026-07-05T08:49:41.842394+00:00"}