Pith. sign in

Paper Citation Record · LEDGER

What do Reward Models Memorize?

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.24484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.24484 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T13:41:43.189796Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3b8e438-53d9-4952-9eaf-08021161d7d0 · outbound

This paper cites The greater this value, the more extreme a features impact on average.

What do Reward Models Memorize? The greater this value, the more extreme a features impact on average

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:43.183116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:43.183116Z digest=sha256:94fa662172d9ea43bc11e1c2cfbbacc12e796387d99463307e2f8ea379fdabaf

Observation 550204db-b0f4-429f-9a5d-84b58c0c080c · outbound

This paper cites an unresolved cited work.

What do Reward Models Memorize? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:43.186446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:43.186446Z digest=sha256:028715964f5ea98e611aa25de9459bf1005c78e3c10885f9d4602c36e0d27589

Observation d89ec3c4-1558-4f11-a917-4702b2fa51bc · outbound

This paper cites InProceed- ings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 18647–18664, Vienna, Austria.

What do Reward Models Memorize? InProceed- ings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 18647–18664, Vienna, Austria

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:41.808291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:41.808291Z digest=sha256:19c7421f36353d22f87bb3c05a17feceead7fc8fa2d8fbc69ff6a861892d1981

Observation 5fecea50-10d8-4ae8-b216-812f8a83f6e4 · outbound

This paper cites InICML 2025 Workshop on Collaborative and Federated Agentic Workflows.

What do Reward Models Memorize? InICML 2025 Workshop on Collaborative and Federated Agentic Workflows

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:41.912434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:41.912434Z digest=sha256:bfcf200cc66642b7d79313a329812657f9b216b5705c155e5d6d9bdfd2e9555d

Observation 2da47083-87dd-4a94-a959-abcdffae1e08 · outbound

This paper cites InThe Thirty-ninth 9 Annual Conference on Neural Information Process- ing Systems.

What do Reward Models Memorize? InThe Thirty-ninth 9 Annual Conference on Neural Information Process- ing Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.062185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.062185Z digest=sha256:633a419e8c12d30e83980e0e0194de2e0d62ebd6a3e5e881195d33b537b9924f

Observation a0a18f50-3a5c-4dde-91a9-4e4c6a427eec · outbound

This paper cites Impact of Fine-Tuning Methods on Memorization in Large Language Models.

What do Reward Models Memorize? Impact of Fine-Tuning Methods on Memorization in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.393478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.393478Z digest=sha256:2c9fc5f202787a258fd9882d913a2b1ba1cb9fc6d7dd5320521b4a47c9317981

Observation 32c34279-9c81-4ede-bb2c-9f550f714cff · outbound

This paper cites InFindings of the Associ- ation for Computational Linguistics: EMNLP 2024, pages 12891–12907, Miami, Florida, USA.

What do Reward Models Memorize? InFindings of the Associ- ation for Computational Linguistics: EMNLP 2024, pages 12891–12907, Miami, Florida, USA

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.610163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.610163Z digest=sha256:7ef1196161ea83433452d3bcab641ed49ca45e03e17dfbfaac7da40714e3c3b1

Observation 2d223efb-27c7-4eba-b131-46810dd16bfd · outbound

This paper cites Proximal Policy Optimization Algorithms.

What do Reward Models Memorize? Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.726088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.726088Z digest=sha256:e963054c4c3ca6d58481b0e74639e2e303e37790163a031b0afd4ee46af54876

Observation 043a5c66-d191-449a-afca-960a7c2be8ca · outbound

This paper cites Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment.

What do Reward Models Memorize? Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.895479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.895479Z digest=sha256:f869a91e019b5150ab29340bdbe3e44432f9dde6e319d61b153d5113f08c73c0

Observation ae456a65-c765-4297-ad50-3c6aa884b0ed · outbound

This paper cites This model was specifically trained on human-LLM interactions Rule-based Rewards {0,1} 16 We annotate responses for possessing one of the identified LLM behaviors in Mu et al.

What do Reward Models Memorize? This model was specifically trained on human-LLM interactions Rule-based Rewards {0,1} 16 We annotate responses for possessing one of the identified LLM behaviors in Mu et al

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:43.125252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:43.125252Z digest=sha256:5977ab2d79b0c4967d21c497cd51e9c4c8ee22ffaeb63a7acc0c7a5665abbc67

Observation e16b773b-8a36-48a3-a21f-9270ac99e8b1 · outbound

This paper cites Unlike the previous metrics, which capture the impact of a feature absolutely, this metric is relative to all other features active in a preference pair.

What do Reward Models Memorize? Unlike the previous metrics, which capture the impact of a feature absolutely, this metric is relative to all other features active in a preference pair

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-07-31T13:41:43.189796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:43.189796Z digest=sha256:d47bd4573a617ee92c7613ae5a1237b955da639dd5d3eaffe899e4b1c19dbd91

Observation 1514c4d5-2b78-40e1-ab1c-695e424d4c9c · outbound

This paper cites 16 0 1 2 3 0.0 0.5 1.0 1.5 Loss 0 1 2 3 0 10 20 30 40 Grad Norm Figure 8: Loss and gradient norm curves for all 25 model iterations when training onPRISM.

What do Reward Models Memorize? 16 0 1 2 3 0.0 0.5 1.0 1.5 Loss 0 1 2 3 0 10 20 30 40 Grad Norm Figure 8: Loss and gradient norm curves for all 25 model iterations when training onPRISM

Reference 1400

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.962985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.962985Z digest=sha256:cbb4e20fff948de652cb6d21573707e0e2788fdd16ae4d3e54f322180b3b672b

Observation 1589a814-1230-4d0d-a3e6-e6e0514900d8 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

What do Reward Models Memorize? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.277767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.277767Z digest=sha256:d9b08149003130bf16d0f83cae2d3a12d3c9cdcafcb3a71f610c91534e5635b2

Observation e271d6b8-d375-436b-b23e-59f13dfa1513 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

What do Reward Models Memorize? A General Language Assistant as a Laboratory for Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:41.701623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:41.701623Z digest=sha256:61e2350c0dd2656b898abbd3861be5ab13bcd08da8f55e09d19e87636d0e6d3b

Observation e0fd8dad-e47a-4213-b904-349fe34cbe70 · outbound

This paper cites InPro- ceedings of the 2022 Conference on Empirical Meth- ods in Natural Language Processing, pages 1816– 1826, Abu Dhabi, United Arab Emirates.

What do Reward Models Memorize? InPro- ceedings of the 2022 Conference on Empirical Meth- ods in Natural Language Processing, pages 1816– 1826, Abu Dhabi, United Arab Emirates

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.495503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.495503Z digest=sha256:bb339b03b4ab309292dd0a63b65db9701a50f410c288b87b5caf7473c8d5ecf2

Observation 8f344aab-7505-4da7-b766-1c0804a95f65 · outbound

This paper cites InProceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 8323–8343, Singapore.

What do Reward Models Memorize? InProceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 8323–8343, Singapore

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.160391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.160391Z digest=sha256:b28c54070c3b677b5b7b05495872f416bd1d2dc4adcf16065efa02e007e80937

Observation 0fa6ef38-b4d4-4936-937a-80e179737dcc · outbound

This paper cites BatchTopK Sparse Autoencoders.

What do Reward Models Memorize? BatchTopK Sparse Autoencoders

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:41.992934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:41.992934Z digest=sha256:809a6b58b1622cc9e1e81130d319b68ddff42548373a1dc935c3891e079bcda1

Observation dfff48d2-40c6-4bc4-aea6-717ae7fd865d · outbound

This paper cites an unresolved cited work.

What do Reward Models Memorize? Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:41.502883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:41.502883Z digest=sha256:2ea52a7863abdb15794916cc63f6128fefd17d2800122ce00dbef1052bcb2bbe

Observation 7662c23f-a6af-492a-9cce-5579b624fca4 · outbound

This paper cites Intel® Corporation.

What do Reward Models Memorize? Intel® Corporation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:41.611493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:41.611493Z digest=sha256:2fcbc8afaa26a0907d63fe464062ce0641af2edafb89b3ea6bb275c0e51778f9

Observation 7813fcd5-2593-44be-8d87-0810b7e3ed01 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

What do Reward Models Memorize? Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 9471

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.807753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.807753Z digest=sha256:21ff5b7927003888c94ba1e8b333734986981b4611359b8597f89f550d5bc1fd

Pith citing papers

No inbound Pith citation observations are available.