Pith. sign in

Paper Citation Record · LEDGER

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning

As of 18 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.19408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19408 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:15:19.895698Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1612773-4807-4786-bf15-1d21f04b9e95 · outbound

This paper cites Dense reward for free in reinforcement learning from human feedback.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Dense reward for free in reinforcement learning from human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.155018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.155018Z digest=sha256:51db98e01ed1ea8e5ea29ff7ec8a20a4ee0e058b7a5db47655127173481a0ad8

Observation 40d49530-339c-4270-b019-946119d79488 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.226461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.226461Z digest=sha256:e0a4789226ada8d7f59eb7a15c6e610626956df78a67543ee864ec68f6ad53ad

Observation 074ee4f9-2342-4368-a7fd-4c482f5cae8a · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.305019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.305019Z digest=sha256:09b310ffcfaa52359cd8c5d2a6878ab64d2a27801065671df0b5eb5aa63a5951

Observation 6d1654b2-3948-476e-b468-3357d054e2fa · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.428563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.428563Z digest=sha256:512f3f0b3c10a14be2d0e44fd9187d54c51276e613babb86eccc887e67ba807b

Observation 46f5c39f-8313-4aa0-9ac9-69f2cf5d1d19 · outbound

This paper cites To- ward semantics-based answer pinpointing.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning To- ward semantics-based answer pinpointing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.490631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.490631Z digest=sha256:b91224c7f30ae68369320a0dbf40861a55f8c9545ce6634179e2ee95d3fb8148

Observation bc50babb-2fbf-4087-9e68-88def37e6412 · outbound

This paper cites Learning question classifiers.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Learning question classifiers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.522302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.522302Z digest=sha256:9479db1253d871cab2c7af2ac2f0b6b51823baa8486eb328edf105cf73d9258f

Observation a0dd0eec-1ae5-4e4b-9bbe-33b6c4f532e3 · outbound

This paper cites The blessing of dimensionality in llm fine-tuning: A variance-curvature perspective,.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning The blessing of dimensionality in llm fine-tuning: A variance-curvature perspective,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.589635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.589635Z digest=sha256:8b82923948b5bf973140bb4ba908a045593d8c0e1070529e983a2cf13053e738

Observation a608cb4e-066c-46d2-a96b-0efe913ebcc4 · outbound

This paper cites Let’s verify step by step.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Let’s verify step by step

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.810036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.810036Z digest=sha256:85f13cba2bebabbe8f4c51d1a869a7ebd909cb30f367fb12505e03b4f2de65d2

Observation 2ebc1bba-b51e-4233-a9f0-d6d44130e546 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.844834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.844834Z digest=sha256:3f4ffca395c2c80581b8badc089d50f2d6891267cd6576f49f0f8640ac266ac6

Observation 95156e19-0ec0-4a9b-8fed-b26b49ecb5f0 · outbound

This paper cites Lee, Danqi Chen, and Sanjeev Arora.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Lee, Danqi Chen, and Sanjeev Arora

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.849014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.849014Z digest=sha256:830953bb658934209786e1b86965df330e03e4d5e2a3af8e5d1c59d18ea508b0

Observation 3b3dc49c-f664-4945-bf09-e5979f0b8983 · outbound

This paper cites Random gradient-free minimization of convex func- tions.Foundations of Computational Mathematics, 17(2):527–566, 2017.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Random gradient-free minimization of convex func- tions.Foundations of Computational Mathematics, 17(2):527–566, 2017

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.852682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.852682Z digest=sha256:57cc8b008c8a02af0200f440aac64da4fa79d6411981bceea69941a84117d780

Observation 1fa9f782-49b7-4896-a4c0-e3fef6f00255 · outbound

This paper cites Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.856120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.856120Z digest=sha256:5173d57c4ffc1b86022d23d2c5887eac39bbab33279260f562480078a9cc42de

Observation a1adf360-e185-4376-8a4b-a3b016651d6b · outbound

This paper cites Qwen2.5 Technical Report.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Qwen2.5 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.860627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.860627Z digest=sha256:0d1f2508ac2beb92413e7207446355465fef07c40327b8a34ceac4b8f3d136d9

Observation cf62eb38-1832-4f7b-8297-a56ac2031945 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.864221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.864221Z digest=sha256:01f3397d7542b95ad8ff09f58709f1da4d271200a2a92bc778223142211d2597

Observation 48289a40-839e-4a98-92ad-a85a1f47eab6 · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Manning, Andrew Ng, and Christopher Potts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.867837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.867837Z digest=sha256:e69b291efd5d4400a49a85792e0131c3865d6991e26734429a13d5260eb37bfd

Observation 850ab672-14f4-424d-8163-5446b0ebc093 · outbound

This paper cites Black-box tun- ing for language-model-as-a-service.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Black-box tun- ing for language-model-as-a-service

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.871408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.871408Z digest=sha256:af42476ef2dffc0e18e5f219a7959928e9a832f4d5a64a2564fcb68d6eed2fd9

Observation 4629c3be-53c4-4be3-892d-318b66d519d9 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.874508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.874508Z digest=sha256:e4d19b381265f97704d6c7ca93ee92137ba055d302eb848e46ecba788c815ca4

Observation 2eb6ba90-93c3-447c-bb2f-ee3a8d42ce8d · outbound

This paper cites TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 18

Resolution
verified exact
doi, observed 2026-08-02T08:18:25.912187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-02T08:15:19.878269Z digest=sha256:277457bc88725dd281009a004b9f066a326eaa8201fd1fef4211d38a63f0eef8

Observation 0beb7bd1-65e7-4566-adbb-47e7684e278a · outbound

This paper cites Opt: Open pre-trained transformer language models, 2022.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Opt: Open pre-trained transformer language models, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.881783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.881783Z digest=sha256:99a12a8db486b303be6678fa8dc87ce668fe9d4ac320b6ff6d7460306190aca4

Observation 10f428bd-cdf0-4a13-b936-84da42b94da2 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.885366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.885366Z digest=sha256:6d08004dd9afa91f45771bb7b195edd8d3409270779acc607cc5ad83d1d95b35

Observation 47647dd0-8559-4301-92c2-4581983fdba1 · outbound

This paper cites Total cost:2KB forward passes.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Total cost:2KB forward passes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.888832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.888832Z digest=sha256:d1962bdf149e79c0ccb4d8ecbc04bc9278e2f9434761f7f7fba781360d7a9a6b

Observation 8c6a4a03-9ce0-4561-b582-cdd0774c74d4 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.892116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.892116Z digest=sha256:c740424abc785a8e4900e8d303ef4df6781b2ccc588874cd1c3b9f4a4fc87411

Observation 096c7ef5-696a-43fc-b7c3-5fe9d76ac254 · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.895698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.895698Z digest=sha256:64236a4df28789528d4e60f2554ccab011f5dc8fe9e6612621dd2ee4545cc7aa

Observation 4a631a5b-30bc-42c8-bba2-f4afe7a0d02b · outbound

This paper cites an unresolved cited work.

Reward-Aware Population Scaling of Evolutionary Strategies in LLM Fine-Tuning Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T08:15:19.670505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:15:19.670505Z digest=sha256:f018eefd333d2f31031af4a1cac9fc763cb2d4feb096f9f42ca204d71cc35816

Pith citing papers

No inbound Pith citation observations are available.