Pith. sign in

Paper Citation Record · LEDGER

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2608.01743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01743 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.368038Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df80e58d-a917-48ca-be6d-d8d867e8e48b · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.245730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.017940Z digest=sha256:86ed43eede3c51df503ce872fa75285a2f12fddd11ca5c7cd55ac57a4d7a30a6

Observation c8b89b9c-8b80-427b-bf58-57a1510d7f77 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.230777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.050129Z digest=sha256:8499b48e54166d5e315e4aa240a6cb5b0d9a0d561a4bf1f8a82b681991421622

Observation f20e3a33-f3ed-47aa-af3c-01ea0516b4d9 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.217308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.060067Z digest=sha256:43070f78b13889746b71aa4d503f09596abcd120a6f72d63b3e75816bd0f5951

Observation d4b0c03d-f000-49f3-a49b-7cb98d1adb80 · outbound

This paper cites H.; Gonzalez, J.; Zhang, H.; and Stoica, I.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning H.; Gonzalez, J.; Zhang, H.; and Stoica, I

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.202641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.069509Z digest=sha256:b0c94e2aae98fdfae6e962ee4a86bee77dac3ad5531dbf9a9af3f047448f80ec

Observation d9356796-39dd-404c-85a1-c7e61bc5c2d1 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.188389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.097267Z digest=sha256:f15c74547fc7d73f03a8619f36b0c3ebd1668b2cb830f3520e40dc863d69568b

Observation 65f5e0f5-bb4a-4284-a5a1-c16010cdfcc8 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.129231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.129231Z digest=sha256:993ae04a059b23ba6b6503a6f850af3c13e4be21d73584e4b4ad072332c970e4

Observation 74af9aad-919c-4603-a4f7-4284782058ba · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.137374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.137374Z digest=sha256:d81ce114a3f719e2b4287333c0768bf1e6d306af004c2bdc9a1f081a00bdd282

Observation 7b0d8687-957a-413b-b072-9ba612e3c9c7 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.161193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.141974Z digest=sha256:b08971d4b92071176ce1b54c228ecee11943bac5dd0d20f3b0478068a6e33bbe

Observation 959a227d-d93c-4771-a638-169caa91aef7 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.146636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.145654Z digest=sha256:f1b0608bca0fb0202777c57be1e3dbf0a3e6abafe4632acdd2e49689d52d2043

Observation 495d56fd-0403-4cdb-a2c7-d82029c5ba61 · outbound

This paper cites an unresolved cited work.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-04T21:54:45.132103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.174384Z digest=sha256:2c0e14cf1157c987308b1f891a373b03e418b8fa809e2b7f288d466154e613f0

Observation 4c92224e-9ff6-4b52-b7dc-a0a876477ea1 · outbound

This paper cites 2024 , journal =.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning 2024 , journal =

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.187684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.187684Z digest=sha256:4a5fe9866e37b5f4949aec2f1b32462f20d677ac97855d9acc41e6f8f04766f4

Observation 09e37af4-78c6-4be7-863b-ec28a04e474f · outbound

This paper cites Proceedings of the 29th symposium on operating systems principles , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Proceedings of the 29th symposium on operating systems principles , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.192272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.192272Z digest=sha256:0423f51db9df8b38f921b759cde2e5bf64e7594ff5702fdffb7b2632a35540d1

Observation f05b61a8-05ed-4c07-954b-c723093d7e39 · outbound

This paper cites Qwen3 Technical Report.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Qwen3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.196474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.196474Z digest=sha256:b9749022ac34c8a33f9ac052240407a94c588da44d28edb7622ac27314bd878f

Observation 81afb144-f8cc-4f4a-be3e-bc8db4b9f015 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.200389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.200389Z digest=sha256:213ec957096ccf4aa4da3d6a41e6b4226e0ed42131ffba522fd5f0e0a4bf2b14

Observation f77e03be-e80f-4784-8276-969b8fc49b18 · outbound

This paper cites Advances in neural information processing systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in neural information processing systems , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.204431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.204431Z digest=sha256:d4737a5b4be03f5dd9a1693b84e5bdfa8cc60d77310824cc73cd818c58d396d7

Observation d4614ed2-ffee-4162-a9b8-5f7d2d541dc7 · outbound

This paper cites Advances in neural information processing systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in neural information processing systems , volume=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.208940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.208940Z digest=sha256:51e53514070751d2600fad900da9a681cf5c9bb9c18182b95d4711a6d0c6fc9a

Observation c5d7d783-687b-4adf-a1dd-938cddac8947 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.213111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.213111Z digest=sha256:6e377fedaac1a0f2c3c521869ff4de6003502ac593137e6c8535cbe60f5363d7

Observation 2aa3844b-7052-4b9d-9a1e-1610e23bdac9 · outbound

This paper cites arXiv preprint arXiv:2509.07430 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2509.07430 , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.217118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.217118Z digest=sha256:3faa1b039557283cbdde72c97ec9bcebd9affe6bbadfb938da578f811581165d

Observation a98d7c5b-c8f8-4eba-b5bd-07ff69b01084 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.221531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.221531Z digest=sha256:b7c73e1dabd496835392b5bd78303cb7effd68c38a90e180f72fe2d69daef2d1

Observation 2a521534-ce38-4f85-9d5e-09031face8e8 · outbound

This paper cites Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.514401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.225531Z digest=sha256:44ae1924398379c07349e8f8c2ad83ce1b7e54c33b120ad3f4e23e8022550746

Observation 4fea440c-0f4f-4aa6-9af2-c869ce602c86 · outbound

This paper cites arXiv preprint arXiv:2510.03865 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2510.03865 , year=

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-08-04T21:54:44.888734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.229256Z digest=sha256:cab091b643d531ae1143dbd5e6f8e2d3509228c68919c179604dc79d61419413

Observation 51920178-91f2-4f57-aa29-e7275570cb94 · outbound

This paper cites SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.782159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.233162Z digest=sha256:00218932f254d862ae7c1856aa9e3b1ce0bdf2db2bca258f01a223ba1f5a8b8d

Observation c5c3b979-95d4-4927-86f1-31f82e92fb49 · outbound

This paper cites arXiv preprint arXiv:2510.20817 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2510.20817 , year=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.237573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.237573Z digest=sha256:05fb4d4eff664a98b31384adfb276f5a2ce13571a9242cf787cee8746705792e

Observation a819ad68-2a25-48bb-ba76-c9fe1bf79dfd · outbound

This paper cites expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.715269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.241343Z digest=sha256:c804bbd3c212ce6085406fcad96344842d6892b9027ac0f9922f656d498d6c6e

Observation 69b0009d-c43c-4ea1-8c06-e629fe386bee · outbound

This paper cites Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.245197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.245197Z digest=sha256:0b1abe84b09628afc6dacea4fb9a2f871037c073faee141451f460301d97aafe

Observation 51f81fca-0199-4d8e-8479-56237d36f1c1 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.248752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.248752Z digest=sha256:c00c5ee505e85f4ab4b30e8e21c8f11cffdc1f2d33d603d31c09965368a5b319

Observation 4b9db30a-cc20-4107-8274-997583ac6163 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.069692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.252587Z digest=sha256:396d3df8556352637dfdf20a6ea73b6b648eaf1f7a8ed8cc7fed679ba30b461d

Observation cb170a3c-c32c-4233-8b8a-7deb11a1ecc8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.256399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.256399Z digest=sha256:b7de872e8091c0e6b602d82112ed66db27c7d15369a7fb666981c32b856d9415

Observation 0008add1-0006-4660-b26c-238db2c742fd · outbound

This paper cites Group Sequence Policy Optimization.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Group Sequence Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.260001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.260001Z digest=sha256:5df29d56706499ff338787432f6fd9e6f9ad22ace6ef852e3e4a7d0825bf2fb5

Observation 0b978fe9-44cc-42d7-996e-074b439cce26 · outbound

This paper cites SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.263730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.263730Z digest=sha256:cfe48cdfc5d01e6ba6bef9b9381a66c83282aa70ccccc6271eab3a7103bd52ce

Observation 61c442a8-efe1-4f94-811d-0f65aa07c352 · outbound

This paper cites arXiv preprint arXiv:2509.21826 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2509.21826 , year=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.267496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.267496Z digest=sha256:0b2ba9632fb4992b07b054831a01ca130c5314ea0897d3ac421cb7d52890d33a

Observation beb7572e-2aad-4fb6-9fb2-415e538b6b4d · outbound

This paper cites ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.693998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.271066Z digest=sha256:4fc1770de63526948b8443db6ee52727c40fe0a44f9f2c6b1c609dad9d257c36

Observation 1760ee1a-2a76-4fac-b3b8-d602362fe35f · outbound

This paper cites Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.275053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.275053Z digest=sha256:de224f8805d54153a2ffaffcccecabb912bae6f95c8ff1e3ca70b0e0eb9f5fa8

Observation 11a28c1b-8692-4df2-bc0f-b48642f39907 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.055465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.279008Z digest=sha256:163a7788884a1e4b7235306cbeb17e0804f42cebca717cc50d61aac2b4f410c7

Observation 513eb1db-d952-4e0c-abc8-c1e864933d55 · outbound

This paper cites arXiv preprint arXiv:2507.14783 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2507.14783 , year=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.282831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.282831Z digest=sha256:d6e4f34c58660c309387847fb5a1bf871139c66311bd90aa08a3bedcc6d9e58f

Observation b0b85819-9a32-450d-990c-f7710a6cf889 · outbound

This paper cites Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.286875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.286875Z digest=sha256:6eddb1f3c891303d050b13ac0d5d142fc32f635ecd132da58295b7a0bfbc17e4

Observation e48110fa-5932-4e42-a981-6716065cab6a · outbound

This paper cites arXiv preprint arXiv:2602.12566 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2602.12566 , year=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.290891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.290891Z digest=sha256:c050e98bc251d5f65b7beef710811bcaa019ac2eff39299a9edf27729a05e0e7

Observation 78f8ee3c-a707-4ca5-881b-389d9ef6b591 · outbound

This paper cites arXiv preprint arXiv:2602.02301 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2602.02301 , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.294669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.294669Z digest=sha256:630e61373582c4d992f703c561489d3d4703e8b3fcbfa3af307a4cfcbf717bc1

Observation 2de762b3-6690-4c75-a29c-ff5df8769829 · outbound

This paper cites arXiv preprint arXiv:2505.17508 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2505.17508 , year=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.298515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.298515Z digest=sha256:49b2ec6e11dbd5ff544b1bf2d725d3f65a47bd8d04d1d6331323584e3f7c254b

Observation f0ab0c6d-c983-4837-ade3-88399ac72e09 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.041362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.302583Z digest=sha256:fada3801afda2aa2257c6348f03e7b7d2af8b364aa50c70940decdfeafa11027

Observation a25a14df-98a3-4de8-be31-241cdd07debb · outbound

This paper cites MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.306632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.306632Z digest=sha256:2f61cbacb962d89a024f99780c148d4fa3d49b99df34b2a8a65d812e4c218c2b

Observation e66e8458-af75-4c5b-8783-51bc09ccd8b2 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:45.025570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.310926Z digest=sha256:941a223026159e3b70d89590b4790d4879290e3f8f1d8bebd6542f5e0a666fad

Observation a5b07d83-b46e-4da2-af82-47e44e99f983 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.315385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.315385Z digest=sha256:556bf7abb9239db582c4ab5ef77f3aa958a68bdb47d0f88acdf6005be8e0258d

Observation fefc93de-90e7-4a81-a7e9-de149a2a8cef · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.319080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.319080Z digest=sha256:d3ce107b9ec7e6dd1364adc3778d75c919f260f45ca306b712582fd4b3f3e8f0

Observation 440a5857-0cef-47f0-9e65-01ef9da28d69 · outbound

This paper cites WildChat: 1M Chat.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning WildChat: 1M Chat

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.322742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.322742Z digest=sha256:5fefec25483e4593dacf6fe166babbe457fd73bc904f7c395e52b65aee48f90a

Observation d85754cd-3a0f-47a7-96bc-de263b245193 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Advances in Neural Information Processing Systems , volume=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.326437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.326437Z digest=sha256:8a0d490a5b17a57050f2604de80b0f4a93c50aeffaa02ccd02eb6bc0cc666642

Observation 6d3c600e-de81-43fc-b9b7-4ad1eebf5f11 · outbound

This paper cites arXiv preprint arXiv:2512.15489 , year =.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2512.15489 , year =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.330165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.330165Z digest=sha256:e198481db00ce4df83d3f2ee1d69b892712495fcc28ba78758434668f1989d56

Observation 5b129d2e-ab42-4199-8d96-22397aa477a2 · outbound

This paper cites arXiv preprint arXiv:2509.20357 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2509.20357 , year=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.333909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.333909Z digest=sha256:3a7284745b894fb0323e9a578e5c3082b335e3a82706f2282b50f2598582d01b

Observation 9f41ce78-a012-42b9-ab64-9f1f61dd0740 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.337692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.337692Z digest=sha256:72f5c60c03fdb717c7fbd6ae027e3b96c498ff7a786d774e44631c6c0bcf4c61

Observation bbd38b4d-117c-43b4-81ca-2cbdc185adab · outbound

This paper cites arXiv preprint arXiv:2512.05962 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2512.05962 , year=

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-08-04T21:54:44.806705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.341274Z digest=sha256:7d64ec0ef4ec5d1a02b9432dcd7b0967cc6a8a3f83201806cd8cde726459d6f1

Observation 490492f5-7b91-4f81-b8db-d1e281bbd6c6 · outbound

This paper cites Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T21:54:44.626662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.345264Z digest=sha256:fbf08b88a62e0bf83100332673e4a22b1c8c9863ea50a37ef22edd56fed980b9

Observation 14284321-4222-44e3-864a-f7944fb9fa81 · outbound

This paper cites arXiv preprint arXiv:2602.19895 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2602.19895 , year=

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.348703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.348703Z digest=sha256:054bceb1225480efe5a58250c4305d685b5ddc26caeda4e5a568ad804533798f

Observation d40d7bb1-73a0-4e5b-804e-e4a11b18e0cf · outbound

This paper cites Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.352386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.352386Z digest=sha256:9cfcac781905e972251c2c0fa3bba98cb8c416e3333d3e82f51c0708bcefc555

Observation 0309255f-307c-41e5-8890-66f275465eba · outbound

This paper cites International Conference on Learning Representations , volume=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning International Conference on Learning Representations , volume=

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:54:44.966856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T21:54:44.356061Z digest=sha256:092e40e776bf4a39c8f167e3493906bfffc2c560c5e015d4ee9abf6b8c8cfbb2

Observation ee7e51f3-7ccf-4a53-b5e3-504e3936633f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.359793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.359793Z digest=sha256:27429f77984ec901226e8285c71bfc58e4242ba3dee5a9a07a3e61b212844eff

Observation 567c4997-dfc7-44bc-90ee-c1b8dd195bdc · outbound

This paper cites arXiv preprint arXiv:2507.14843 , year=.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning arXiv preprint arXiv:2507.14843 , year=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.363782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.363782Z digest=sha256:a1beb99c47a70f4af593b70fe1abe822ffaa084c2d02e43117085330b24f8e05

Observation 4a43e7f0-9217-49ef-97ad-685026069a89 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.368038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.368038Z digest=sha256:8194288eb9e3e2c9c21eb80f8281ad26388c5c54026358e04f29188411d28a64

Pith citing papers

No inbound Pith citation observations are available.