{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:RRELWZBVJQT6BMEJSSWZ3327PP","short_pith_number":"pith:RRELWZBV","schema_version":"1.0","canonical_sha256":"8c48bb64354c27e0b08994ad9def5f7bc744b34059cebe8dd92e6767a8629a18","source":{"kind":"arxiv","id":"2509.24372","version":3},"attestation_state":"computed","paper":{"title":"Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.NE"],"primary_cat":"cs.LG","authors_text":"Babak Hodjat, Conor F. Hayes, Elliot Meyerson, Qiyao Liang, Risto Miikkulainen, Roberto Dailey, Xin Qiu, Yinggan Xu, Yulu Gan","submitted_at":"2025-09-29T07:19:34Z","abstract_excerpt":"Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has emerged as the dominant fine-tuning paradigm, underpinning many state-of-the-art LLMs. In contrast, evolution strategies (ES) has largely been overlooked due to the widespread belief that it does not scale to modern model sizes. This paper overturns this assumption by demonstrating the first successful application of ES to full-parameter fine-tuning of LLMs at the billion-parameter scale, without dimensionality reduction. ES can indeed search over extreme"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2509.24372","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.LG","submitted_at":"2025-09-29T07:19:34Z","cross_cats_sorted":["cs.AI","cs.NE"],"title_canon_sha256":"f52c128c1eb4a1d1f31f3293dd084560cf7cbd33685b90ae408f56ce18d4a573","abstract_canon_sha256":"f810b0802e788a30b396824688447dd3e30f1ae6915a722eac38f33c80449ccf"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-15T01:21:46.927878Z","signature_b64":"KxAFFsoGZUbwHjNJGnM0clZjc6cwDQ8fP1rSRAz9E0EgmuGuuDZBRCv0uXOef6IAT2WG/VeA9QkWCDu/MHREAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"8c48bb64354c27e0b08994ad9def5f7bc744b34059cebe8dd92e6767a8629a18","last_reissued_at":"2026-07-15T01:21:46.926977Z","signature_status":"signed_v1","first_computed_at":"2026-07-15T01:21:46.926977Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI","cs.NE"],"primary_cat":"cs.LG","authors_text":"Babak Hodjat, Conor F. Hayes, Elliot Meyerson, Qiyao Liang, Risto Miikkulainen, Roberto Dailey, Xin Qiu, Yinggan Xu, Yulu Gan","submitted_at":"2025-09-29T07:19:34Z","abstract_excerpt":"Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has emerged as the dominant fine-tuning paradigm, underpinning many state-of-the-art LLMs. In contrast, evolution strategies (ES) has largely been overlooked due to the widespread belief that it does not scale to modern model sizes. This paper overturns this assumption by demonstrating the first successful application of ES to full-parameter fine-tuning of LLMs at the billion-parameter scale, without dimensionality reduction. ES can indeed search over extreme"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2509.24372","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2509.24372/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2509.24372","created_at":"2026-07-15T01:21:46.927400+00:00"},{"alias_kind":"arxiv_version","alias_value":"2509.24372v3","created_at":"2026-07-15T01:21:46.927400+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2509.24372","created_at":"2026-07-15T01:21:46.927400+00:00"},{"alias_kind":"pith_short_12","alias_value":"RRELWZBVJQT6","created_at":"2026-07-15T01:21:46.927400+00:00"},{"alias_kind":"pith_short_16","alias_value":"RRELWZBVJQT6BMEJ","created_at":"2026-07-15T01:21:46.927400+00:00"},{"alias_kind":"pith_short_8","alias_value":"RRELWZBV","created_at":"2026-07-15T01:21:46.927400+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":5,"sample":[{"citing_arxiv_id":"2606.12279","citing_title":"Mathematical perspective on genetic algorithms with optimization guided operators","ref_index":40,"is_internal_anchor":true},{"citing_arxiv_id":"2606.30619","citing_title":"Why can genetic algorithms work in high-dimensional search spaces?","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2605.16345","citing_title":"Goal-Conditioned Supervised Learning for LLM Fine-Tuning","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2605.16727","citing_title":"PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play","ref_index":54,"is_internal_anchor":true},{"citing_arxiv_id":"2602.01003","citing_title":"ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning","ref_index":5,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP","json":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP.json","graph_json":"https://pith.science/api/pith-number/RRELWZBVJQT6BMEJSSWZ3327PP/graph.json","events_json":"https://pith.science/api/pith-number/RRELWZBVJQT6BMEJSSWZ3327PP/events.json","paper":"https://pith.science/paper/RRELWZBV"},"agent_actions":{"view_html":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP","download_json":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP.json","view_paper":"https://pith.science/paper/RRELWZBV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2509.24372&json=true","fetch_graph":"https://pith.science/api/pith-number/RRELWZBVJQT6BMEJSSWZ3327PP/graph.json","fetch_events":"https://pith.science/api/pith-number/RRELWZBVJQT6BMEJSSWZ3327PP/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP/action/timestamp_anchor","attest_storage":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP/action/storage_attestation","attest_author":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP/action/author_attestation","sign_citation":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP/action/citation_signature","submit_replication":"https://pith.science/pith/RRELWZBVJQT6BMEJSSWZ3327PP/action/replication_record"}},"created_at":"2026-07-15T01:21:46.927400+00:00","updated_at":"2026-07-15T01:21:46.927400+00:00"}