Pith. sign in

Paper Citation Record · LEDGER

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 8 inbound Pith citation observations for arXiv:2509.24372.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.24372 v3

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:43:29.945895Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:06.694233Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

83 of 83 outbound references displayed

  • verified exact7
  • verified fuzzy0
  • unresolved76
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ffed6561-48be-4931-a9be-ae1b5c170ceb · outbound

This paper cites GPT-4 Technical Report.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:16.944824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:16.944824Z digest=sha256:7d24073f42bd22e31dda36c53267a71485b459c49f25812a7910d347dbab3f0b

Observation 4a2b85df-7f2a-42e6-a1c6-84f78c69bb68 · outbound

This paper cites Intrinsic dimensionality explains the effectiveness of language model fine-tuning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Intrinsic dimensionality explains the effectiveness of language model fine-tuning

Reference 2

Resolution
verified exact
doi, observed 2026-08-04T14:43:27.589291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T14:43:17.104745Z digest=sha256:a170d499cd439888e76a7b6bd2ae057dfc121113950bdb2c9d0ac915000e2e66

Observation d92eb794-c4cf-461a-887c-578877622de4 · outbound

This paper cites Llama 3 model card, 2024.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Llama 3 model card, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.244748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.244748Z digest=sha256:74ea30ed189accc565c9c95e91fa439fd815ca234c9cc318f7dbaf6066cb24de

Observation 70bbafff-d02b-452a-989e-25dfb20a7668 · outbound

This paper cites Evolutionary optimization of model merging recipes.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolutionary optimization of model merging recipes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.443331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.443331Z digest=sha256:927deabb20974bb71080b8b0a5e6cfe03577e89a728d497c28becd5ca5cec5cb

Observation 1cf56819-0c16-469d-a314-258c68a9f53b · outbound

This paper cites Introducing Claude 4, 2025.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Introducing Claude 4, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.695291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.695291Z digest=sha256:1f0172a6d5688650246ee17952cc41e7d5c251c98c0348d6fd87693c704f3fe4

Observation a775b486-7838-497b-bbfe-946235b2934d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:17.934749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:17.934749Z digest=sha256:d86d9b9ae497dac7ccdc155a5ca2566380e48ac5f5cb7f44e910d8c1fd4aad19

Observation debbf38d-33ca-4e10-aaa2-9c6cfd9ecb91 · outbound

This paper cites Understanding pre-training and fine-tuning from loss landscape perspectives.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Understanding pre-training and fine-tuning from loss landscape perspectives

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.060517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.060517Z digest=sha256:e8f0b0498a79842bbaf4a6f906a38a72122d27d88a184dee75b7eeab2069d342

Observation 40e0d99b-048c-4be0-af24-9cac8f899fea · outbound

This paper cites On the weaknesses of reinforcement learning for neural machine translation.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning On the weaknesses of reinforcement learning for neural machine translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.304745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.304745Z digest=sha256:425f5d09063cf0bd05f4fb3d42cfa7e7d09af4d20f13236c2dc8feb5130c2ea0

Observation 1448cc4e-5477-4bc3-98e4-0d327e955730 · outbound

This paper cites Back to basics: benchmarking canonical evolution strategies for playing atari.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Back to basics: benchmarking canonical evolution strategies for playing atari

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.515288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.515288Z digest=sha256:279384fe4a6d6b80584851cb2cf21234239e815ecca7c97eb237416d07fb3eca

Observation ccda3c47-12d6-4009-a746-794a5dffbac2 · outbound

This paper cites Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.684764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.684764Z digest=sha256:531edae2ee011bf941ab0f106ed47b9cece16fb37cf21ec5e5359d164d180968

Observation b1cadf89-8871-4d75-9cd2-ec44aed1fb11 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:18.815282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:18.815282Z digest=sha256:b05f21f6cc1b0230e081fe1875d620484868203f3393972b981adb36f730b21b

Observation 116f2248-7d3d-4cfb-b0cd-7c9ff0d9f382 · outbound

This paper cites Knowledge fusion by evolving weights of language models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Knowledge fusion by evolving weights of language models

Reference 12

Resolution
verified exact
doi, observed 2026-08-04T14:43:27.258406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T14:43:18.912503Z digest=sha256:fa51127bc4f851dffc75ceb312e942328837983ed0545a2c3a94a041fcac7c12

Observation b97881ce-5a75-415f-b8b7-f4646c3872b1 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Detecting hallucinations in large language models using semantic entropy

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.144753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.144753Z digest=sha256:9360bf3e2fa35b59b58b9d1b9080248cebf594aefcd88c84e1904b91a2f876c2

Observation 97dabea9-1298-426d-8912-3e57462f0dc7 · outbound

This paper cites Reward shaping to mitigate reward hacking in RLHF.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Reward shaping to mitigate reward hacking in RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.404754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.404754Z digest=sha256:ebcd821556974edd5569dfe948b3c8306a37b5d05887c34ffae244dc05091278

Observation ef754ffa-0924-4677-9409-758a7b7867d4 · outbound

This paper cites Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective ST ars.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective ST ars

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.604788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.604788Z digest=sha256:fc17009bc30a57efd0c3cd9620e1a6c378e849be6e61b3a3819f0da622c91987

Observation 1cbe0367-b8b7-4903-a123-a226527c892b · outbound

This paper cites Scaling laws for reward model overoptimization.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Scaling laws for reward model overoptimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:19.857855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:19.857855Z digest=sha256:3da45028a472a052ae304442561095fda2bce3e29a3ca9a039fd22f6f3d726f2

Observation 289e40bd-8b2f-428e-afed-01bc37e65ae0 · outbound

This paper cites Deep learning, volume 1.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deep learning, volume 1

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.014401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.014401Z digest=sha256:c4038c6c41f02f2a1115b344919ebdef945cd79b15482951baee22d0d995e146

Observation 2178f773-8149-4329-991e-b487a03a513f · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., 2025.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities., 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.204798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.204798Z digest=sha256:7e0275f818422faa6510e47d46c58c5f107857c6d2c14d97382545208e705669

Observation d62c2a83-2353-42b8-a766-2c584ced21d5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.390576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.390576Z digest=sha256:1de263f2fdffc4c89d3a99c6246b82c9c9330748b503253e49d0aa76615aa551

Observation a7036ecf-177d-481a-a876-e664d29aaf9b · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:20.764759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:20.764759Z digest=sha256:ebe0c49f85cb15b818352273f49b79e1eae51407b1acb9853d8c80f5fa2946cf

Observation ea0ae0a5-8475-4dfd-a79d-d2757b03ddaf · outbound

This paper cites Connecting large language models with evolutionary algorithms yields powerful prompt optimizers.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Connecting large language models with evolutionary algorithms yields powerful prompt optimizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.051034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.051034Z digest=sha256:abb9484e0c9f803352006ee66de20796c6fc339c2261999f53927a356afe7d67

Observation 1dba326f-0240-4c76-9f03-f8b5a2b7d07f · outbound

This paper cites Completely derandomized self-adaptation in evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Completely derandomized self-adaptation in evolution strategies

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.314845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.314845Z digest=sha256:e018c378822891fd7ddd234b0c168e099f185248aa615b0647b91699548aa6a5

Observation 18b5cfba-bd1e-4db9-b35f-19b29e9e4410 · outbound

This paper cites When evolution strategy meets language models tuning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning When evolution strategy meets language models tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.479943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.479943Z digest=sha256:77f4e173ea3fc58b508c423cbef06515f200ac6c3c6d83fcc9e8cc5860dc7873

Observation 9e173aea-9f86-4c5d-9ce1-329aff70b79b · outbound

This paper cites Neuroevolution for reinforcement learning using evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Neuroevolution for reinforcement learning using evolution strategies

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.755222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.755222Z digest=sha256:bc9f1438a459ec672014060c66fc91dbf212f69892573257ecb0bfbe669397db

Observation 66a764d1-4096-4b16-ac11-46cd21524b4a · outbound

This paper cites Do we need to verify step by step? rethinking process supervision from a theoretical perspective.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Do we need to verify step by step? rethinking process supervision from a theoretical perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:21.904752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:21.904752Z digest=sha256:2332b033f754f7f59e6c321198b835e8b59891e85a03d65d5a2b3a70083c265c

Observation ec6ea8c0-8122-4c0d-ab22-977bbb2993b4 · outbound

This paper cites Mixtral of Experts.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Mixtral of Experts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.124747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.124747Z digest=sha256:787e9b709474a33415f0366ad6711f401d623062b4cb21743e8044deda285f63

Observation b00a1de9-0509-45d4-888f-42b39fc7ed0e · outbound

This paper cites Derivative-free optimization for low-rank adaptation in large language models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Derivative-free optimization for low-rank adaptation in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.334749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.334749Z digest=sha256:3141d51fd2a3d5f683fc0cf5957aed032b4f244b8a24cab81592b0defcd92ca7

Observation 2bb53075-2a59-4fca-b33f-64d302e00a34 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.504761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.504761Z digest=sha256:5ac59253fe26d6863fa9da081b8b38fd526a324e2973b864b748f535c7ffb7c1

Observation ecbb5529-880d-4d57-98c6-d33f02e7a269 · outbound

This paper cites Fine-tuning chatgpt for automatic scoring.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning chatgpt for automatic scoring

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.674753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.674753Z digest=sha256:c30229c5b3410c938187663047ae5a65979d5a1efe017d05f74fb33ebd756e0c

Observation 26c35849-849e-4738-9715-b2b353ed2dc9 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:22.920603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:22.920603Z digest=sha256:518e634593a9a3005344602f475b9a4d182afa61fe2286cb9f7f21866aeebe7e

Observation 955ec27f-6a47-4a9f-bf3d-64b382b1f973 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.084747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.084747Z digest=sha256:3d3986a2a6a11c29d595c82431d7132faf47e93d7a88584b5fb76caa0b7f16b8

Observation 3ee762dd-3f1a-4278-b0a8-5ac1b622d931 · outbound

This paper cites DeepSeek-V3 Technical Report.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeek-V3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.187732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.187732Z digest=sha256:29ff8c4759429ff3c653d55bbe11107ee1ff8c6dc6b803aa51536d7c762c01bb

Observation 71fa5a7d-e6e4-49a7-8ff6-bf5646c24d07 · outbound

This paper cites Sparse me ZO : Less parameters for better performance in zeroth-order LLM fine-tuning, 2025.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sparse me ZO : Less parameters for better performance in zeroth-order LLM fine-tuning, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.344787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.344787Z digest=sha256:37c6a9a98635b98181dc68c5cca59e6d97aea3f9d46386f8e1427a15b9dd4eee

Observation efa07711-73e6-4449-971f-3fdcc02cff0b · outbound

This paper cites Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.484928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.484928Z digest=sha256:b0011af1d43df34be9000bd754ee715b2dc6e4a3466988eaaa94c8d3d5c7b5fc

Observation c5df2399-1981-44ba-968f-11fd4aadbd6c · outbound

This paper cites Fine-tuning language models with just forward passes.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning language models with just forward passes

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.574795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.574795Z digest=sha256:f4ad5f6d636d14a222b38e01c8dc4c1f67a66db26c582cdd464a54a5b1d985d0

Observation a1cfbdc1-a4a6-490a-95fc-77c111fe4117 · outbound

This paper cites Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Nelson, Herbie Bradley, Adam Gaier, Arash Moradi, Amy K

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.664754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.664754Z digest=sha256:678c605e1ff5056ea8ab2e4e1d3891c2aec6ccaf4090fefff77aaea8757d7860

Observation 9788c601-c4b5-4d58-b29e-254452e3d465 · outbound

This paper cites What is artificial superintelligence?, 2023.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning What is artificial superintelligence?, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.764751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.764751Z digest=sha256:ef373cb3d2c8eb6f44d498bc02e0aad3b85f9ca505001f2223da55c79d3d830d

Observation 3bb4a197-ca86-4646-abab-89e9b6fe6c97 · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:23.915163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:23.915163Z digest=sha256:fe27eb92b4a1b6e10975155de2471b8fe293c0e5717b1b80da55ace91e5426dd

Observation 9c197115-645a-4f2e-acd0-9db7f5d71f97 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.294757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.294757Z digest=sha256:791a12d462bc6c6508e22b51e04945d6f166ba397bd9654e2e7b063e3c7b945c

Observation fd170c12-df61-4bbd-a660-c882d507609f · outbound

This paper cites Tinyzero.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Tinyzero

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.415276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.415276Z digest=sha256:1f13817686d140d960d4e56018df445269ec76323f4eafa0385a7d0f30216809

Observation 49155d3b-e903-42c3-a3cc-b920057a2491 · outbound

This paper cites Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.624754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.624754Z digest=sha256:03856a5f0930dce721cab32acac2982114c476bff30377a172744159ddb60b82

Observation 49a2d506-5899-4534-a9bc-98b445aefd8f · outbound

This paper cites Semantic density: Uncertainty quantification for large language models through confidence measurement in semantic space.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Semantic density: Uncertainty quantification for large language models through confidence measurement in semantic space

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.787960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.787960Z digest=sha256:dad264820dfb3c81f15c300a9bd362c9c19f02073af23be1e78bd9b6631b496c

Observation 8e7782ce-fd85-4227-b42a-3ca5eb9fa557 · outbound

This paper cites Manning, and Chelsea Finn.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Manning, and Chelsea Finn

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:24.934759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:24.934759Z digest=sha256:b77ab3b20746afdb53d004152fad31d28df65304103d41521d16952bb34cb763

Observation 9b669db8-8f19-4a24-b1d5-b8806868d090 · outbound

This paper cites Rechenberg.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Rechenberg

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.124738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.124738Z digest=sha256:85700ce28a1d3fc20d07a791e0c8583201f27d8686b79978203e71c4756b20cc

Observation 4ae9bac0-f017-4e12-b2f2-e81f9066e6fe · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.304763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.304763Z digest=sha256:59ef3a5f058bb7e80bebe94c272b08e59ae32c0af86d7374db58ac99c2504b66

Observation ce995643-c078-4e58-9949-c8c8c866b226 · outbound

This paper cites Pawan Kumar, Emilien Dupont, Francisco J.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Pawan Kumar, Emilien Dupont, Francisco J

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.494752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.494752Z digest=sha256:5778d31d39ea0938b8fdd0a686653f918f07634436e19ccc13f1b96b8593486f

Observation 4167e780-7675-4cd8-8104-dc36299a2d36 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Code Llama: Open Foundation Models for Code

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.564745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.564745Z digest=sha256:5893a55496940249c406e5cf7241b18a37c3ae10adeae4571bccec71850a8111

Observation cd28f1ad-0bf8-4a59-bcbb-b5d4477a3d64 · outbound

This paper cites u ckstie , Martin Felder, and J \.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning u ckstie , Martin Felder, and J \

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.646284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.646284Z digest=sha256:40c9f32ebc983a788bf06bc9ec6fadd93bae215525dbc96303d96c557d851306

Observation 0563fd3f-7e40-414d-9928-66ac14c64292 · outbound

This paper cites u ckstie , Frank Sehnke, Tom Schaul, Daan Wierstra, Yi Sun, and J \.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning u ckstie , Frank Sehnke, Tom Schaul, Daan Wierstra, Yi Sun, and J \

Reference 49

Resolution
verified exact
doi, observed 2026-08-04T14:49:22.811963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T14:43:25.755143Z digest=sha256:82f7242dca00c98d3c94173b3a8b3006581893ae5b015f15b764ea4176865f12

Observation 2aa166e2-0e31-4c69-9a1c-63de63bc6595 · outbound

This paper cites Improved techniques for training gans.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Improved techniques for training gans

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.834738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.834738Z digest=sha256:0e8259cc4ac1fd53989754c0a12996d53c6b888995fec812545d054e54a86b5d

Observation a1127278-1185-41d4-b12a-b84bf56d8fc6 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:25.910796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:25.910796Z digest=sha256:77e337bede5b7d00c47463ad773a6d3201d1ebbe007a08fbf5bf14fe8148569f

Observation 1c269cfe-a434-47ab-b6c9-81da6b750074 · outbound

This paper cites How well can a genetic algorithm fine-tune transformer encoders? a first approach.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning How well can a genetic algorithm fine-tune transformer encoders? a first approach

Reference 52

Resolution
verified exact
doi, observed 2026-08-04T14:49:22.554438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T14:43:26.187680Z digest=sha256:31a2cdb8b7d84046574dde4f8ea3f0584991b18159d61d453bce3f1040deb7c5

Observation 56aaed94-1d56-4d14-9a96-75d295c6917c · outbound

This paper cites Approximating kl divergence, 2020.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Approximating kl divergence, 2020

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.317534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.317534Z digest=sha256:5475c87707e45011202766d8a6460d85ca37c384da4c1668ff2d1caab8e75a68

Observation 1a2957ed-851b-410d-9d65-94988b818858 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.444749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.444749Z digest=sha256:638a99632a1f165119030bc30bddb707384c946fbd85640f958bb56afea1e6a3

Observation 566bc3db-07ed-4778-b751-b2526b58273f · outbound

This paper cites Numerische Optimierung von Computermodellen mittels der Evo-lutionsstrategie, volume 26.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Numerische Optimierung von Computermodellen mittels der Evo-lutionsstrategie, volume 26

Reference 55

Resolution
verified exact
doi, observed 2026-08-04T14:49:22.266846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T14:43:26.595147Z digest=sha256:59336de349b24a5bf285d5c62e96ad40d354571daa52e128f2edfa47770d1d78

Observation 021a19c9-2ada-4bc5-ad18-5b943bbf7621 · outbound

This paper cites Parameter-exploring policy gradients.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Parameter-exploring policy gradients

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.744734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.744734Z digest=sha256:0284600c39ea4320829971e39fdf45b17f419e5c61f85d39467d524e13ec4497

Observation 5299c176-c970-41f7-96bf-b9b272b65ff1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.852975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.852975Z digest=sha256:e41fc9f7f233f8160f5a8b4d6224914c553b67040a21395a4134a868596c970a

Observation 80b23c43-5d60-420b-8886-818a98ec6671 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Towards Expert-Level Medical Question Answering with Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:26.960792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:26.960792Z digest=sha256:b0878ee29ae7ade982be4150bf304f201905742760509da04a2ef31bbcc00944

Observation f9c38cc8-5575-47ec-ae15-97c78348dc49 · outbound

This paper cites PRMB ench: A fine-grained and challenging benchmark for process-level reward models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning PRMB ench: A fine-grained and challenging benchmark for process-level reward models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.025445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.025445Z digest=sha256:9689fa260370fd9fc9e6ad0ef4d547d740f4ba59c429e537bf81b821bc04664d

Observation 557e0319-cb47-450b-958b-5bc0fad95dd4 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.126234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.126234Z digest=sha256:b62c4dc06db56990d9e5bdb098e792195843d6f119a8bd139642cc6077a6a999

Observation 73b5aced-ca17-4bc6-8dc1-f4374deccddc · outbound

This paper cites A Technical Survey of Reinforcement Learning Techniques for Large Language Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning A Technical Survey of Reinforcement Learning Techniques for Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.174129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.174129Z digest=sha256:a99f1d9b9efa1448f8728f438cbbdf244b490ce156cb46961edd25b7d8c24fba

Observation 0991185e-9785-45f6-88f2-3b8b01088df3 · outbound

This paper cites Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.246511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.246511Z digest=sha256:18b7fa596095d9bd8a751d0fad29604ac07870ee21d734392ffa0e577bef9159

Observation e3c6fe67-92a7-4f31-a208-e7efb0061d44 · outbound

This paper cites BBT v2: Towards a gradient-free future with large language models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning BBT v2: Towards a gradient-free future with large language models

Reference 63

Resolution
verified exact
doi, observed 2026-08-04T14:49:21.997293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T14:43:27.374747Z digest=sha256:9e7050400dfc8eaf3bf6282bbdf1003a71d1d95b615d124dc3eaacdf6c344eaa

Observation b60689ed-5695-4a7d-9fbc-0ae90570da75 · outbound

This paper cites Black-box tuning for language-model-as-a-service.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Black-box tuning for language-model-as-a-service

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.448579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.448579Z digest=sha256:2a1b1022d6db8b98c87fe419cc897b6ec5983a72aab9e1c0bc417dda1f984ee4

Observation 250e12ca-d12d-4cac-8249-0cb6d5168f45 · outbound

This paper cites Sutton and Andrew G.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Sutton and Andrew G

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.506331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.506331Z digest=sha256:e516e466eb2949e591865857be9a63310704ace97c21aac751d60fd1e5a61b27

Observation 3e7e83d6-a8de-44e9-8cab-319f86980e50 · outbound

This paper cites Fine-tuning mt5-based transformer via cma-es for sentiment analysis.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Fine-tuning mt5-based transformer via cma-es for sentiment analysis

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.586741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.586741Z digest=sha256:70bb5ba16c9d32bce8ed9086eb01a862a6c238bdad64f13073b5bd70b908ef24

Observation 93e0ef8b-6fba-4725-ab17-0d57ec27dc60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.686155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.686155Z digest=sha256:6546334da377b08dbba031e51c2ebac4ad8eb4ae72b8ea9a2a71aec347869e31

Observation 8fd0b78e-4c31-44bd-9bf2-edcd19b001ad · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Solving math word problems with process- and outcome-based feedback

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:27.886374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:27.886374Z digest=sha256:e711f15e6ed6c5db67a61f8d3928ccf2bb351450fdf55485acc2d0d583fa5851

Observation 2139c937-4a60-4a80-89cd-74df5b9b7a4a · outbound

This paper cites Andrew Bagnell.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Andrew Bagnell

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.045066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.045066Z digest=sha256:063c90bc7679586c124f5ba4ba4bb8c5138b442954ab5f75148b9cf0d2b3aad3

Observation 4ae3ebca-2647-47f9-8859-f6d4bec14fb3 · outbound

This paper cites When large language models meet evolutionary algorithms: Potential enhancements and challenges.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning When large language models meet evolutionary algorithms: Potential enhancements and challenges

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.108064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.108064Z digest=sha256:1f76a0818422806c2683d0f8f8473c64913617cd7576bb1eed661eac462910b8

Observation 03e3f036-81ac-425e-a968-a3352427da62 · outbound

This paper cites Natural evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Natural evolution strategies

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.295794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.295794Z digest=sha256:0272ad103218323f74b244e913d1df7cb52717f0b4f22f888af079a607268603

Observation 20c6296a-1bc3-4eed-b3bb-27e7ec4859ad · outbound

This paper cites Natural evolution strategies.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Natural evolution strategies

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.396359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.396359Z digest=sha256:db3b5e33a0e47b5d6df6bf4c32cbf5e17b1e9c2c08f8b370e5aead01513e63d9

Observation fa1b9614-df6d-42fb-a2b3-d94f7d0fad23 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning BloombergGPT: A Large Language Model for Finance

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.534733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.534733Z digest=sha256:678cb2bf7849ea211a3de5994a202c162b1e2d57df0b5e9755839278578bb880

Observation d2d6254c-a23b-4b63-bddf-c1b91fc01fee · outbound

This paper cites Evolutionary computation in the era of large language model: Survey and roadmap.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Evolutionary computation in the era of large language model: Survey and roadmap

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.654757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.654757Z digest=sha256:69ec715cd2757aa72fb26cb6aac96cb4b09db21433de36116a66f2b1053b3f0d

Observation 09d98d04-fe78-4afc-86f1-89a15c6fd4f5 · outbound

This paper cites Qwen2.5-1M Technical Report.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Qwen2.5-1M Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.822086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.822086Z digest=sha256:1cd19dff530d457c3634db0d59c1e4754b1177ef3c0d17471d288fa7f45e98cc

Observation 4e249551-1814-407a-93e2-139675d2d875 · outbound

This paper cites On the Relationship Between the OpenAI Evolution Strategy and Stochastic Gradient Descent.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning On the Relationship Between the OpenAI Evolution Strategy and Stochastic Gradient Descent

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:28.925179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:28.925179Z digest=sha256:3ae30b5918e4cc3c957b6a049bbff50d71e0ddebb680c2db410c4c71a869bb5b

Observation c862fcb1-73f3-4966-b6bd-508ff12f6e4e · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.015865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.015865Z digest=sha256:c559acb1e2960bac4d678797fd86cc803e44bd9b879090acdb963fb5d345f1cd

Observation 7e12e69a-8c2d-48c9-aea7-2f66a762bb1d · outbound

This paper cites Genetic prompt search via exploiting language model probabilities.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Genetic prompt search via exploiting language model probabilities

Reference 78

Resolution
verified exact
doi, observed 2026-08-04T14:49:21.859507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-04T14:43:29.112085Z digest=sha256:7e6a2d9b3ada1d634427d5adc216d66619d8d7788d41d7252eb9d6e02051c618

Observation cea816d8-261a-442c-b0f6-df389bed24ac · outbound

This paper cites DPO meets PPO : Reinforced token optimization for RLHF.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning DPO meets PPO : Reinforced token optimization for RLHF

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.366209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.366209Z digest=sha256:22e91a958b6f064846bf5cebfbd68a95fa08abd42189f17ab0b669e30c42fd87

Observation a6d9c6a0-4abc-4204-8649-b723d2a9f70d · outbound

This paper cites write newline.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning write newline

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.468661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.468661Z digest=sha256:b0fd79dbcc5c0d7941dd8d1f7a103f6c8baa4269f6dae3f7a598957de2061fa1

Observation 3140f06a-fd27-456b-be01-aaf1d282cf6d · outbound

This paper cites @esa (Ref.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning @esa (Ref

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.624735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.624735Z digest=sha256:6d57cfe70feb1dbcf60670232cef4147d0e740df11def0272d394c1ab4ab37c4

Observation 96b62c46-9c97-4f38-a72e-af0097894cdf · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.761120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.761120Z digest=sha256:c83e56ec0a0bf22e51eae0b4acf329ae8b245cc82df04ac86baafdc3025c36e0

Observation b6c09985-6d31-4ba6-aefe-be9ccd556db6 · outbound

This paper cites an unresolved cited work.

Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:29.945895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:29.945895Z digest=sha256:90865b7c9c385aa862517c8f9dc6385f7b7fa66ef76485912da49fc8ec5bce71

Pith citing papers

Observation 08cfb055-de24-4a25-92bc-b230408cc6d1 · inbound

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training cites this paper.

Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:54:47.861802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:54:47.861802Z digest=sha256:70b61f06483d4cbb924a46d46b3d35eec4a51856626a8de92a68c93951f2490f

Observation d7d77328-35f2-43e4-a15a-13a73d071b8f · inbound

ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning cites this paper.

ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:24:04.873255Z digest=sha256:e52dd7275abab3963673dbc26d54265dfe29aac35c95b939b1ddb82a0a964c39

Observation c079d1a8-4bcd-4cf5-b57b-8adc4393a860 · inbound

Goal-Conditioned Supervised Learning for LLM Fine-Tuning cites this paper.

Goal-Conditioned Supervised Learning for LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:37:46.345159Z digest=sha256:0c24f51646b576cd4baf7c5e368715f895f9cb9bf890823c15e69364423603f0

Observation 4dbd30cf-5d48-4e6f-9b4f-3bcbc479737f · inbound

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play cites this paper.

PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T21:37:56.570173Z digest=sha256:92b64619ac984bae6a18a41dac978ccc50ead92aa821c1cd62f4e93448cdef17

Observation 124a61b3-fdbf-48b9-b17a-ef78c1143029 · inbound

Mathematical perspective on genetic algorithms with optimization guided operators cites this paper.

Mathematical perspective on genetic algorithms with optimization guided operators Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T07:34:35.422699Z digest=sha256:7246f2c036e0c4af67b7452bee4bd53a1a146ee65a6099b0a247ff1553c3a3e1

Observation 56bb7a7c-7b55-4bf2-ba20-fa69215bdf9f · inbound

Why can genetic algorithms work in high-dimensional search spaces? cites this paper.

Why can genetic algorithms work in high-dimensional search spaces? Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-15T02:21:57.501664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T03:11:04.083187Z digest=sha256:9f49dafbed08c8f166b9a429762edac2b9eb2232b0832acbfefc86c8734a6d5f

Observation 1fa9f782-49b7-4896-a4c0-e3fef6f00255 · inbound

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning cites this paper.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.856120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.856120Z digest=sha256:77c97af9956de7a02e7ec3e8f76c9cfda81a02d9109bc1a4baa5dfdef66ad1ed

Observation a9361b78-916d-42e4-ab8c-c08c8c56e984 · inbound

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging cites this paper.

Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:06.694233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:15:06.694233Z digest=sha256:a24c2835e7436e40bfc2023e008d0b478bdf04e80ba9f428738de4b2dcc274de