Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T13:41:43.189796Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.24484.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T13:41:43.189796Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c3b8e438-53d9-4952-9eaf-08021161d7d0 · outbound
What do Reward Models Memorize? The greater this value, the more extreme a features impact on average
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550204db-b0f4-429f-9a5d-84b58c0c080c · outbound
What do Reward Models Memorize? Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89ec3c4-1558-4f11-a917-4702b2fa51bc · outbound
What do Reward Models Memorize? InProceed- ings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 18647–18664, Vienna, Austria
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fecea50-10d8-4ae8-b216-812f8a83f6e4 · outbound
What do Reward Models Memorize? InICML 2025 Workshop on Collaborative and Federated Agentic Workflows
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da47083-87dd-4a94-a959-abcdffae1e08 · outbound
What do Reward Models Memorize? InThe Thirty-ninth 9 Annual Conference on Neural Information Process- ing Systems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0a18f50-3a5c-4dde-91a9-4e4c6a427eec · outbound
What do Reward Models Memorize? Impact of Fine-Tuning Methods on Memorization in Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c34279-9c81-4ede-bb2c-9f550f714cff · outbound
What do Reward Models Memorize? InFindings of the Associ- ation for Computational Linguistics: EMNLP 2024, pages 12891–12907, Miami, Florida, USA
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d223efb-27c7-4eba-b131-46810dd16bfd · outbound
What do Reward Models Memorize? Proximal Policy Optimization Algorithms
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043a5c66-d191-449a-afca-960a7c2be8ca · outbound
What do Reward Models Memorize? Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae456a65-c765-4297-ad50-3c6aa884b0ed · outbound
What do Reward Models Memorize? This model was specifically trained on human-LLM interactions Rule-based Rewards {0,1} 16 We annotate responses for possessing one of the identified LLM behaviors in Mu et al
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16b773b-8a36-48a3-a21f-9270ac99e8b1 · outbound
What do Reward Models Memorize? Unlike the previous metrics, which capture the impact of a feature absolutely, this metric is relative to all other features active in a preference pair
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1514c4d5-2b78-40e1-ab1c-695e424d4c9c · outbound
What do Reward Models Memorize? 16 0 1 2 3 0.0 0.5 1.0 1.5 Loss 0 1 2 3 0 10 20 30 40 Grad Norm Figure 8: Loss and gradient norm curves for all 25 model iterations when training onPRISM
Reference 1400
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1589a814-1230-4d0d-a3e6-e6e0514900d8 · outbound
What do Reward Models Memorize? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e271d6b8-d375-436b-b23e-59f13dfa1513 · outbound
What do Reward Models Memorize? A General Language Assistant as a Laboratory for Alignment
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0fd8dad-e47a-4213-b904-349fe34cbe70 · outbound
What do Reward Models Memorize? InPro- ceedings of the 2022 Conference on Empirical Meth- ods in Natural Language Processing, pages 1816– 1826, Abu Dhabi, United Arab Emirates
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f344aab-7505-4da7-b766-1c0804a95f65 · outbound
What do Reward Models Memorize? InProceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 8323–8343, Singapore
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa6ef38-b4d4-4936-937a-80e179737dcc · outbound
What do Reward Models Memorize? BatchTopK Sparse Autoencoders
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfff48d2-40c6-4bc4-aea6-717ae7fd865d · outbound
What do Reward Models Memorize? Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7662c23f-a6af-492a-9cce-5579b624fca4 · outbound
What do Reward Models Memorize? Intel® Corporation
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7813fcd5-2593-44be-8d87-0810b7e3ed01 · outbound
What do Reward Models Memorize? Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 9471
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.