Pith. sign in

Paper Citation Record · LEDGER

Robust Policy Optimization to Prevent Catastrophic Forgetting

As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 3 inbound Pith citation observations for arXiv:2602.08813.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08813 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T05:33:42.965249Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T23:08:30.205397Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T23:09:12.140809Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact45
  • verified fuzzy6
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 711248ff-1ee5-4c7f-a6ba-5ef3038ffab1 · outbound

This paper cites GPT-4 Technical Report.

Robust Policy Optimization to Prevent Catastrophic Forgetting GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.357484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:933d2bd3c183eb671f12f80a4df9e8e1abeee5b8ccb1f6b5d287c3316ad91b54

Observation 559d6437-d8e9-4d57-adcd-dd29305c1b6c · outbound

This paper cites Better Fine-Tuning by Reducing Representational Collapse.

Robust Policy Optimization to Prevent Catastrophic Forgetting Better Fine-Tuning by Reducing Representational Collapse

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.401469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:039554e8f59e9a81620979e48af958e1c4ee5999e83b2cd8216dfb69abd08b04

Observation fa6aefeb-02bf-462d-9177-0c714affd228 · outbound

This paper cites OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs.

Robust Policy Optimization to Prevent Catastrophic Forgetting OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.407770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:9d931131f3ed6acda79d5041003c592696c7ecc6b954f8f888d5f39eeb582ae2

Observation 86e69fdf-0b75-4915-8ccb-57bb03930e1d · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Robust Policy Optimization to Prevent Catastrophic Forgetting Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.453179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:2db92de877f240b5a5e63596eb3c8dbf2d11401a0fc4c880035d586cde6df012

Observation f0d798f2-4573-4d8c-af0c-9b7f52e60b90 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Robust Policy Optimization to Prevent Catastrophic Forgetting Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.330637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:558ffa1370a41f7de3500464e3168a9d239faaae5b7fc89045981ec03d1b727f

Observation df6a0323-ab72-4f7d-ab46-11d34c64f68b · outbound

This paper cites Program Synthesis with Large Language Models.

Robust Policy Optimization to Prevent Catastrophic Forgetting Program Synthesis with Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.341644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:31fb31e6d505f073b829be1900659ae6c0d46bc78629b9f2bf60e4fb94ca723f

Observation 2af55432-7fac-43f2-9e39-a8e65e071d7e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Robust Policy Optimization to Prevent Catastrophic Forgetting Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.347502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:5064bb8d2a94df5924da19d9e83bcb5079a1f4cdf0a7342750e0bc7ff4b37288

Observation cb46b6ba-5ba2-4736-a388-6ce7d144ba22 · outbound

This paper cites LoRA Learns Less and Forgets Less.

Robust Policy Optimization to Prevent Catastrophic Forgetting LoRA Learns Less and Forgets Less

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.414444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:5a10bb00a1a638c4b80c50e843bdf911ff35e445f30fad7454fbfd38b71e62b2

Observation af71d845-a5aa-4221-943c-dff081acfc1c · outbound

This paper cites Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei.

Robust Policy Optimization to Prevent Catastrophic Forgetting Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:37:24.756687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:65d46b1ac9527584876d81907051397272e79157616d5bc75bf9a54d56de470b

Observation c5cb076e-2b2c-4f28-8600-314feb1b314d · outbound

This paper cites Beyond variance reduction: Understanding the true impact of baselines on policy optimization.

Robust Policy Optimization to Prevent Catastrophic Forgetting Beyond variance reduction: Understanding the true impact of baselines on policy optimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:37:24.760264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:ef0d878d118e0263925dae9466ee9d58e2988d89c58b52df3f602bbd2db2b2c1

Observation c49341f6-3159-469f-88a2-669d8563c5b6 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Robust Policy Optimization to Prevent Catastrophic Forgetting Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.389775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:93dc9d24c1bdef77cac082b3fdf925681fda60458b36684f9b04c9b43a1b0ee0

Observation d64a2717-7ee8-463c-8673-3601c2b96f57 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Robust Policy Optimization to Prevent Catastrophic Forgetting UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:41:29.205174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:eed6840d8f6583eebeac4cc0f2c09ea242bf464843a1d0fae8d89b5ba0da8510

Observation 1c22e9c0-1e31-49c9-85d9-05b0b27cb9f3 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Robust Policy Optimization to Prevent Catastrophic Forgetting OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.313558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:af7bf4b9e3d3344a40d4470d74dab930a38be4027c0fc10fa63cc89c88820bbf

Observation 919c6647-3049-4180-b935-9c4d6966324a · outbound

This paper cites Distributional Robustness and Regularization in Reinforcement Learning.

Robust Policy Optimization to Prevent Catastrophic Forgetting Distributional Robustness and Regularization in Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.433302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:dcef3894ff1ec0fa592ec7aa86bc3e4c42da664892cf1bb5020f7284986a19e5

Observation ac994747-fe2e-4d5e-aed2-ae18e0486d5f · outbound

This paper cites SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging.

Robust Policy Optimization to Prevent Catastrophic Forgetting SafeMERGE: Preserving Safety Alignment in Fine-Tuned Large Language Models via Selective Layer-Wise Model Merging

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.307536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:71ea55f998c890c789e0bbcbc255c5aef76f2021c36a930cd584fffe1c6c4ee8

Observation 7c8a1f8c-512f-4627-a737-1e2586556a52 · outbound

This paper cites Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping.

Robust Policy Optimization to Prevent Catastrophic Forgetting Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.444353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:e3a6200de9955f1c5233d581b30ecdcae29017af2e871e7c9e5eb2bacba48e90

Observation b5e77811-fd14-4ebd-b11c-f058f01278b2 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Robust Policy Optimization to Prevent Catastrophic Forgetting PaLM-E: An Embodied Multimodal Language Model

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.325187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:ee74b2f1a739133ed5513361319660eeeaad784d2dba04b9e09c991a0f05beb6

Observation e2735e6a-ddad-4c94-a5ac-f5f002fc71dd · outbound

This paper cites Maximum Entropy RL (Provably) Solves Some Robust RL Problems.

Robust Policy Optimization to Prevent Catastrophic Forgetting Maximum Entropy RL (Provably) Solves Some Robust RL Problems

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.301959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:82560b917158217432d4e3c03b2b71edb5f03b11981d1df9fc0a0b8da351cc2d

Observation 5c79820f-dec0-4b52-b1d6-250bb116af50 · outbound

This paper cites Sharpness-Aware Minimization for Efficiently Improving Generalization.

Robust Policy Optimization to Prevent Catastrophic Forgetting Sharpness-Aware Minimization for Efficiently Improving Generalization

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:13:53.559812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:227a07620abe398b98ae67263e1f4d879c1b4924754f28c71b4b139ea931fe02

Observation 6703db9a-8f80-42f9-8fc7-14f99051e50f · outbound

This paper cites Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails.

Robust Policy Optimization to Prevent Catastrophic Forgetting Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.278142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:17c85c56965ae6291f646ad394e7c73ee398d82ff5dae6dcf1f26e7c76807dc1

Observation 244123ad-62a2-4b24-a5be-5212803cfefe · outbound

This paper cites An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks.

Robust Policy Optimization to Prevent Catastrophic Forgetting An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.458645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:895d0b020af6b72c0c2cb9473f83e1e4a9ca81b8c3d5d3ff6f62e16712c5c66c

Observation 2f3b6aa5-fc1b-498b-847c-ee1eadbf510c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Robust Policy Optimization to Prevent Catastrophic Forgetting Measuring Mathematical Problem Solving With the MATH Dataset

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.427211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:355cd48f3c9f279fe9e5b2eda9babc418009c9760ff8e6448ef5474b2bff2874

Observation 86865fd9-7ed0-40f1-a556-8d2affb4102d · outbound

This paper cites Re-evaluating Continual Learning Scenarios: A Categorization and Case for Strong Baselines.

Robust Policy Optimization to Prevent Catastrophic Forgetting Re-evaluating Continual Learning Scenarios: A Categorization and Case for Strong Baselines

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.319177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:72a3b1aef2a3c04f87d1e357aa0c1e5fadb4cd5d456dec34fff9937daec485ee

Observation 239e695f-550f-4302-8fa3-892a3eaff3e8 · outbound

This paper cites Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal.

Robust Policy Optimization to Prevent Catastrophic Forgetting Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.336512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:776232940794f636cbeeb8529604a49e0529f24fc14ff6dce071e810739444ea

Observation 177faa3d-e1b7-4869-8a41-e072ab1103c3 · outbound

This paper cites Editing Models with Task Arithmetic.

Robust Policy Optimization to Prevent Catastrophic Forgetting Editing Models with Task Arithmetic

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.378755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:f70377dfbf055cad6e902ef1096b6597783c7a2e9c9b97d65655fc91290c654e

Observation 092973c8-8eae-420d-90bd-bbd81d52d99b · outbound

This paper cites Mistral 7B.

Robust Policy Optimization to Prevent Catastrophic Forgetting Mistral 7B

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T05:37:24.264725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:530d94c3bd142a3a4448bc54f980305a4fc32aeb06befbd698b5ab618346286c

Observation af016d18-9411-470d-9135-74f150fcd6f8 · outbound

This paper cites Fantastic Generalization Measures and Where to Find Them.

Robust Policy Optimization to Prevent Catastrophic Forgetting Fantastic Generalization Measures and Where to Find Them

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.235960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:645670f82679d495c3ba48176a7946456259287c8550c8ab421cab523b7bb4b5

Observation 8c7482f0-606d-44fb-aae7-277a8950dd65 · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

Robust Policy Optimization to Prevent Catastrophic Forgetting On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.295802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:0edd635b432e8e8700dc7b3ca98b8a24a0de108adb2d9a25c91b17c6c88f3082

Observation df779ed7-e923-42f7-b09f-64515349fb03 · outbound

This paper cites Reasoning as an adaptive defense for safety.arXiv preprint arXiv:2507.00971.

Robust Policy Optimization to Prevent Catastrophic Forgetting Reasoning as an adaptive defense for safety.arXiv preprint arXiv:2507.00971

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.271685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:de25557ac12f4f1e7cc475631ba5731fff3e90efe84c82b6d684aa84fe901a3b

Observation bc2752e7-8bb0-4f05-b5a1-da963b26de79 · outbound

This paper cites Understanding Catastrophic Forgetting in Language Models via Implicit Inference.

Robust Policy Optimization to Prevent Catastrophic Forgetting Understanding Catastrophic Forgetting in Language Models via Implicit Inference

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.290928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:917451c0397e4c21b6af164fffa14cd6238d66d5106eb19288c57917be250011

Observation 922364bf-4b1b-47be-b57a-d6dad9d5925e · outbound

This paper cites Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models.

Robust Policy Optimization to Prevent Catastrophic Forgetting Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.211591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:511813fc7e762dde441ee8458ec6c5fffd2bce11391192a5226c343e04130c5f

Observation 6ac63863-d362-4649-b63c-6837705b42aa · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Robust Policy Optimization to Prevent Catastrophic Forgetting Understanding R1-Zero-Like Training: A Critical Perspective

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.448704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:c722cc9dacf7aa5fc0857356d8374913444ba4b4eb531d4fbefd00aac1539962

Observation a27f83c0-0c9b-49f8-b061-3cf31ee5d617 · outbound

This paper cites RewardBench 2: Advancing Reward Model Evaluation.

Robust Policy Optimization to Prevent Catastrophic Forgetting RewardBench 2: Advancing Reward Model Evaluation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.368495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:851a3241b97b145e8023cf6cb96a775539214b9fcdb1d749870677e44f18d056

Observation b3fe66d7-da98-4239-986e-0af177aa631f · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Robust Policy Optimization to Prevent Catastrophic Forgetting HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.242197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:f127260045684203512759c0364a0b9ae384b28ace55643d03151196fa040293

Observation d7530b65-0677-48cb-aa17-087339cb63dc · outbound

This paper cites The Llama 3 Herd of Models.

Robust Policy Optimization to Prevent Catastrophic Forgetting The Llama 3 Herd of Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.253867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:1291734158ad2e93bfc3f40bed48ecb9df76f242b57d028058b29667ce59efc1

Observation 5b49cae1-41cb-4833-ba23-794ffa72a39e · outbound

This paper cites URL http://www.jstor.org/stable/2334280.

Robust Policy Optimization to Prevent Catastrophic Forgetting URL http://www.jstor.org/stable/2334280

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.259485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:5a4df9b8a5c27797c28b9f0e3a402e5208338028b34079138e62dd5f32e9b16b

Observation 7a8e1104-6315-44f5-b626-3e0e40a78f38 · outbound

This paper cites URL https: //doi.org/10.48550/arXiv.2502.02421.

Robust Policy Optimization to Prevent Catastrophic Forgetting URL https: //doi.org/10.48550/arXiv.2502.02421

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.384172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:4297fbd0aea75b97dcdabd56fc70725f6c9462f1fdca798013efe916199979de

Observation 60fc2eed-06da-449b-8601-336ebe4ed3bd · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Robust Policy Optimization to Prevent Catastrophic Forgetting Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.284160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:797112cd441f415340ba72a8bebe281bbdbebdf4ff94328d64b876d301d4f45c

Observation c534c6af-9e58-40e0-bbba-33d68b94a85d · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Robust Policy Optimization to Prevent Catastrophic Forgetting Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.169526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:3817c5c3b32bc92282e248a3c9546fa938b5a4aef9b5632be1e2c39a4ca44b6d

Observation 91151e86-ecd9-4831-ba26-3afcfa3bc548 · outbound

This paper cites URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ b0bc711f48724237b38823c4d9cee10b-Paper-Conference.pdf.

Robust Policy Optimization to Prevent Catastrophic Forgetting URL https://proceedings.neurips.cc/paper_files/paper/2024/file/ b0bc711f48724237b38823c4d9cee10b-Paper-Conference.pdf

Reference 40

Resolution
verified exact
doi, observed 2026-05-16T05:37:23.431910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:d608f22fcc210e0541c9bfa1eeaf043423f86bef09453b16003be49bc729d469

Observation 9cba3312-c59f-443e-ad8b-cee6c8b57291 · outbound

This paper cites Qwen2.5 Technical Report.

Robust Policy Optimization to Prevent Catastrophic Forgetting Qwen2.5 Technical Report

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T05:37:24.395303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:a2f7924ec872ce00afd06255c47c818ce6401334a7d78a0b9b532809be1f7112

Observation a040000d-e1bd-4a67-ba26-a0b2c698085a · outbound

This paper cites Progressive Neural Networks.

Robust Policy Optimization to Prevent Catastrophic Forgetting Progressive Neural Networks

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.229501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:8a2160050dd626ed6eeafa5efd7e4783d70c595615b4248b28f3d1cc5c02d668

Observation 985e76a7-aa02-4e98-b989-56bc3d46651c · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

Robust Policy Optimization to Prevent Catastrophic Forgetting Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.247698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:9cb4f9d39834f16475a5d2ef37308f1b1b1ab2c2e169dc661d60bdbf5a64cc42

Observation 152bf980-b4ad-4f3f-8360-2edc4efc0855 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Robust Policy Optimization to Prevent Catastrophic Forgetting Proximal Policy Optimization Algorithms

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.217120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:b68d6dd749b9d307b20627b0e4ff2b89351021d42c46dba238bd920573e4c12e

Observation 430cc406-25a6-4570-bdbb-387d0dce9743 · outbound

This paper cites Fine-tuned Language Models are Continual Learners.

Robust Policy Optimization to Prevent Catastrophic Forgetting Fine-tuned Language Models are Continual Learners

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.187688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:48348787cff83e95b47838086ff605f77dd68085adf5f56084fc909212e9cc6f

Observation 82f7e0a7-e209-490d-bfe8-95b7e96d8bb5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Robust Policy Optimization to Prevent Catastrophic Forgetting DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.204851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:086b7b0efea860706cbf1c06f1fcc99922ee2c7cb971fb9d7017ecce30327438

Observation 0b4d2595-65fe-4a61-ad00-1e9dd4b490ec · outbound

This paper cites URL https://doi.org/10.1137/16M1058297.

Robust Policy Optimization to Prevent Catastrophic Forgetting URL https://doi.org/10.1137/16M1058297

Reference 47

Resolution
verified exact
doi, observed 2026-05-16T05:37:23.424772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:9456ef7c2b771259b6bd325772d30a6afb58d72f9a60347c81adddc49091c04a

Observation 6438140a-e23d-4e4b-b802-83ede3650685 · outbound

This paper cites Aman Sinha, Hongseok Namkoong, Riccardo V olpi, and John Duchi.

Robust Policy Optimization to Prevent Catastrophic Forgetting Aman Sinha, Hongseok Namkoong, Riccardo V olpi, and John Duchi

Reference 48

Resolution
verified exact
doi, observed 2026-05-16T05:37:23.440495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:653417afc47bc728c2bc1b2001f26c9cfdf187118dc7d5f07d4d589cf9528918

Observation c2d74b44-f3a3-4721-a9c6-bccefb716797 · outbound

This paper cites Distributionally Robust Reinforcement Learning.

Robust Policy Optimization to Prevent Catastrophic Forgetting Distributionally Robust Reinforcement Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:37:24.363375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:9e0f71d307c61ace8963323ce46ad4e14bfd80f5e76de403d79ddf7237911342

Observation 43fe7f04-86a7-4e31-a4a9-1c34dec6fa4b · outbound

This paper cites LAMOL: LAnguage MOdeling for Lifelong Language Learning.

Robust Policy Optimization to Prevent Catastrophic Forgetting LAMOL: LAnguage MOdeling for Lifelong Language Learning

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.180992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:c19edecf21b0ec6b190358027e82a2e423686ce9f7fcfaa4a014859f315881cc

Observation 3cb9b195-6a58-47f2-9ecf-447dd540ddd0 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Robust Policy Optimization to Prevent Catastrophic Forgetting Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.175635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:099d461a6cc2c867b11e0c1facb5f9154406060f1bd0c1cc83fb8ae40aab600e

Observation a2c9efde-983c-4d1b-a6f0-6533b68a3cb2 · outbound

This paper cites Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.

Robust Policy Optimization to Prevent Catastrophic Forgetting Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.193897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:cdf9e1661f5424e5f8986d480c891b39870a7c7ebdfbf919a4101ccfc6f767a8

Observation 5cb617dd-44fd-407f-a252-7838b2959ade · outbound

This paper cites Robust Reinforcement Learning using Adversarial Populations.

Robust Policy Optimization to Prevent Catastrophic Forgetting Robust Reinforcement Learning using Adversarial Populations

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.352963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:22a74f1334984e7701e9f3a6e121776fc365247a74c89dfcc7d694adcaaae2c6

Observation 1422cd46-37eb-416a-b4ef-b2e6a9fe433e · outbound

This paper cites Orthogonal subspace learning for language model continual learning.

Robust Policy Optimization to Prevent Catastrophic Forgetting Orthogonal subspace learning for language model continual learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:37:24.781137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:0ce77c2c9b2fc19d1d8caba22f12805fbfb0c8e70b97978c27307861b42604a3

Observation 30a2df69-b7c7-4aa0-bc5f-0db29504e29f · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Robust Policy Optimization to Prevent Catastrophic Forgetting Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:37:24.200106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:6cdf1e1760bad79981af295842881b6158acef43f9151702ae6cb4b038ef6669

Observation 4f08c312-b145-4507-a47f-db9af96708e9 · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Robust Policy Optimization to Prevent Catastrophic Forgetting Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.223558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:0e5bf0eee7163382f4941c04276ab8646dabf8e74874b99562d8c96c792ad552

Observation c14935d7-8666-48bc-a347-a7b9b035360d · outbound

This paper cites Surrogate Gap Minimization Improves Sharpness-Aware Training.

Robust Policy Optimization to Prevent Catastrophic Forgetting Surrogate Gap Minimization Improves Sharpness-Aware Training

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.439668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:2a9a9b5118da8466923af3525727824baf89691d86a2d24f037fa2a5f441a5fc

Observation 9b24b51c-f914-4b40-b6fd-2abe69296052 · outbound

This paper cites The phenomenon arises because gradient updates for new objectives overwrite parameters critical to earlier tasks (Goodfellow et al., 2013; Kirkpatrick et al., 2017).

Robust Policy Optimization to Prevent Catastrophic Forgetting The phenomenon arises because gradient updates for new objectives overwrite parameters critical to earlier tasks (Goodfellow et al., 2013; Kirkpatrick et al., 2017)

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:37:24.774134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:265884a96fd446647a4d79b156bff89246d6280c4190660d109411ce5f7908b8

Observation aff41f45-2618-469f-8cab-f79b9bb8ae38 · outbound

This paper cites an unresolved cited work.

Robust Policy Optimization to Prevent Catastrophic Forgetting Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-16T05:37:24.770591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:68ba5bd1d4db68c0886b3fb8c710477191dd7efc50ae32b5d62cb643e2ea67b8

Observation 7c85fcd1-618c-4e7c-ade8-d21749acfce2 · outbound

This paper cites an unresolved cited work.

Robust Policy Optimization to Prevent Catastrophic Forgetting Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-16T05:37:24.777281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:a74b577abcf5146d49c9b6973637c3a2c511ba05d79228b654e11e8901a4c2a3

Observation aec5963e-6f6e-4877-8c56-61fed1cabefd · outbound

This paper cites Llama-3.1-8B-Instruct-RM-RB2.

Robust Policy Optimization to Prevent Catastrophic Forgetting Llama-3.1-8B-Instruct-RM-RB2

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:37:24.764654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:aa542615a887967dd40a14caeeff99860181bd804e369c8fa01178982a5d1a5f

Observation d88f0e38-1922-41e1-9efd-b88963fe09ee · outbound

This paper cites \boxed{ }.

Robust Policy Optimization to Prevent Catastrophic Forgetting \boxed{ }

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T05:37:24.767696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:6c036f86aa6e789158368ff0b6eb7924c9fd917b39a8042db335c4346087e757

Observation 7de89430-bdb8-474f-9de3-2bfbccec6b52 · outbound

This paper cites Llama-3.1-8B-Instruct-RM-RB2.

Robust Policy Optimization to Prevent Catastrophic Forgetting Llama-3.1-8B-Instruct-RM-RB2

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T05:37:24.749224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:801ff1f58964056b517e0aa095efcbae7f4723d41430c8d8037874899bfd2b63

Observation 1fc2e0c9-cca5-434b-b0b7-0ba562e269a2 · outbound

This paper cites an unresolved cited work.

Robust Policy Optimization to Prevent Catastrophic Forgetting Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-16T05:37:24.752584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:4cc1a8c6a8d621b2f8b3b3ad112db008cbd10e09421358bb78d8d5dc9b6f1ecf

Pith citing papers

Observation 1a7679e3-73f3-4cc0-9da2-94c35e1b005e · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Robust Policy Optimization to Prevent Catastrophic Forgetting

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:48:15.075895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:d7d2531f92679abd85af16d3bcbf70c2834f1396a043d73d45b222e62d345bf3

Observation 956b72aa-9034-4509-a890-2b3f5121eb85 · inbound

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT cites this paper.

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT Robust Policy Optimization to Prevent Catastrophic Forgetting

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:48:15.075895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:31:33.829939Z digest=sha256:617880e97fb9508608de3b037b8f13bfded3e75999095e098132bae1b7db79dc

Observation 611c45d5-63a8-4178-b300-747ef127bff1 · inbound

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT cites this paper.

Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT Robust Policy Optimization to Prevent Catastrophic Forgetting

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:09:12.143062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T23:08:30.205397Z digest=sha256:0945cb1e4b336774d462d71678e921231ea61b1fe9c236ffe571ecdf1a4536b5