Pith. sign in

Paper Citation Record · LEDGER

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning

As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.19408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19408 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:15:19.895698Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1612773-4807-4786-bf15-1d21f04b9e95 · outbound

This paper cites Dense reward for free in reinforcement learning from human feedback.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Dense reward for free in reinforcement learning from human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.155018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.155018Z digest=sha256:901bfb2880f74f52bc8a6ce1cab4ca38850a3f922ac8313273e26a28bc9c8fe3

Observation 40d49530-339c-4270-b019-946119d79488 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.226461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.226461Z digest=sha256:62642a552b851475a74bcc6acaec9156e238b7f1a6e6916d65e13a1d682c8381

Observation 074ee4f9-2342-4368-a7fd-4c482f5cae8a · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.305019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.305019Z digest=sha256:b0bcec48d3a8281e35de678688835028ee11266f8778d9d7ce140607d053e2ff

Observation 6d1654b2-3948-476e-b468-3357d054e2fa · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.428563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.428563Z digest=sha256:88f08650ef8a236c7fa807fbf2ae521a834efb2a1dbe75eca8f76f22225458df

Observation 46f5c39f-8313-4aa0-9ac9-69f2cf5d1d19 · outbound

This paper cites To- ward semantics-based answer pinpointing.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning To- ward semantics-based answer pinpointing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.490631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.490631Z digest=sha256:90c43215ac0282eaadb110a4c35dd7c5c021150178f48f6bd7d8cf5e52a00dc9

Observation bc50babb-2fbf-4087-9e68-88def37e6412 · outbound

This paper cites Learning question classifiers.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Learning question classifiers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.522302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.522302Z digest=sha256:5ffe8f0a89cc6ead33f2b84794afacb696ad69156a01c8b33252329b3a11156b

Observation a0dd0eec-1ae5-4e4b-9bbe-33b6c4f532e3 · outbound

This paper cites The blessing of dimensionality in llm fine-tuning: A variance-curvature perspective,.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning The blessing of dimensionality in llm fine-tuning: A variance-curvature perspective,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.589635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.589635Z digest=sha256:7b6b1ba304fc4fcd1192a37fe9b819f6b66ed8e8f19bde1a5e2d8b71070d8d13

Observation a608cb4e-066c-46d2-a96b-0efe913ebcc4 · outbound

This paper cites Let’s verify step by step.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Let’s verify step by step

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.810036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.810036Z digest=sha256:1e5d8bc1a375427a1c4730825d38310523e1f41d993e9eb85d6e16754bbd3eac

Observation 2ebc1bba-b51e-4233-a9f0-d6d44130e546 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.844834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.844834Z digest=sha256:782937a9fd59a8f655bd61d4bd7fb9fa4cd94b2faff7b2d477c8d2c01d255032

Observation 95156e19-0ec0-4a9b-8fed-b26b49ecb5f0 · outbound

This paper cites Lee, Danqi Chen, and Sanjeev Arora.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Lee, Danqi Chen, and Sanjeev Arora

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.849014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.849014Z digest=sha256:9ecbcaa9b2508e19a990710fbd61d657d434b8cba6bf426c8198bb744c46d52f

Observation 3b3dc49c-f664-4945-bf09-e5979f0b8983 · outbound

This paper cites Random gradient-free minimization of convex func- tions.Foundations of Computational Mathematics, 17(2):527–566, 2017.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Random gradient-free minimization of convex func- tions.Foundations of Computational Mathematics, 17(2):527–566, 2017

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.852682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.852682Z digest=sha256:8b63cd60fdb8ed31856be4a0f22edc8d343db50bdf0297a7cb27ad3fd55f313c

Observation 1fa9f782-49b7-4896-a4c0-e3fef6f00255 · outbound

This paper cites Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.856120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.856120Z digest=sha256:68806aa3f24b4f3b463f67c89c3a0aa43130128318022e4a0616f2b4a6124c3c

Observation a1adf360-e185-4376-8a4b-a3b016651d6b · outbound

This paper cites Qwen2.5 Technical Report.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Qwen2.5 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.860627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.860627Z digest=sha256:5ec9bb380fabf178e85b075ecdcefba47ff973e5511c2ef82e0cc58e1e593768

Observation cf62eb38-1832-4f7b-8297-a56ac2031945 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.864221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.864221Z digest=sha256:d3e6b3e4816fcc35e5f155dca0d9122539e230fc7e64f5b69e675382fbb9199a

Observation 48289a40-839e-4a98-92ad-a85a1f47eab6 · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Manning, Andrew Ng, and Christopher Potts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.867837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.867837Z digest=sha256:7dd5a0cd5ce99919354306eb40e6fedfcc19ba158d0be5ec659dfa19f36be91e

Observation 850ab672-14f4-424d-8163-5446b0ebc093 · outbound

This paper cites Black-box tun- ing for language-model-as-a-service.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Black-box tun- ing for language-model-as-a-service

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.871408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.871408Z digest=sha256:5759ac5b0a66397dacb090a4a3d79ec4779fb4dfcb60d900c777d9dc997164b9

Observation 4629c3be-53c4-4be3-892d-318b66d519d9 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.874508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.874508Z digest=sha256:b5be617d75ee9304c2bda159eafe7bfd84278f9f2cc89e61bf79acd875c29ed9

Observation 2eb6ba90-93c3-447c-bb2f-ee3a8d42ce8d · outbound

This paper cites TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 18

Resolution
verified exact
doi, observed 2026-08-02T08:18:25.912187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-02T08:15:19.878269Z digest=sha256:43f02e94520af9f16cdc85e9ac59024136661e2b47c813c047d107cd421833a1

Observation 0beb7bd1-65e7-4566-adbb-47e7684e278a · outbound

This paper cites Opt: Open pre-trained transformer language models, 2022.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Opt: Open pre-trained transformer language models, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.881783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.881783Z digest=sha256:247f196b0cc646444f5947fbd3c2d468a104774b9c484811d61c6bf3daa244ca

Observation 10f428bd-cdf0-4a13-b936-84da42b94da2 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.885366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.885366Z digest=sha256:e7aa30ab222bd26566f3be969676a4f8b712b6e2e6569753cbad5656f4152d49

Observation 47647dd0-8559-4301-92c2-4581983fdba1 · outbound

This paper cites Total cost:2KB forward passes.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Total cost:2KB forward passes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.888832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.888832Z digest=sha256:d35b198bfae2d54aff1506d8bd55b035e01f18a9579f42a86ad76ebdc37ff4f0

Observation 8c6a4a03-9ce0-4561-b582-cdd0774c74d4 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.892116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.892116Z digest=sha256:cfebcbb6725900537ec398bfa80f23535da166e07b2a2af9fa1f771893a68e4f

Observation 096c7ef5-696a-43fc-b7c3-5fe9d76ac254 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.895698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.895698Z digest=sha256:7e5e4acc29b98685c1f2fba0a2798f898fa2a314e9783a222c2ab6bdcf41006c

Observation 4a631a5b-30bc-42c8-bba2-f4afe7a0d02b · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.670505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.670505Z digest=sha256:30005ac567c03e82becfdbf3f8672948c346b2e9f7f9a1d1caa3d300c9a50504

Pith citing papers

No inbound Pith citation observations are available.