Pith. sign in

Paper Citation Record · LEDGER

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

As of 4 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 2 inbound Pith citation observations for arXiv:2508.20697.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20697 v3

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T20:40:44.496392Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:47:40.850633Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact30
  • verified fuzzy34
  • unresolved4
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbcf669b-aeaa-4456-b980-02db88afd7a1 · outbound

This paper cites https://aimodelplace.com.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning https://aimodelplace.com

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.816342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:4b627fc2b05729ebbbf7a718b8971bbbf88aaf5b95a186b5781ca1c7e6826797

Observation 33507515-4787-4a21-a406-8575b012b042 · outbound

This paper cites https://azure.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning https://azure

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.804552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:d2eb50fabe998579fc20255c72e6504f5858b1d84d410bf7d0162def1ba8dca3

Observation 8dc8158f-2b44-4c4c-b849-37582792a725 · outbound

This paper cites https://docs.mistral.ai/ guides/finetuning.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning https://docs.mistral.ai/ guides/finetuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.801243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:f824bab4ecaefda607a8bea3563a8b22b38d5a9da2a9a06c422a61c23c2e49e4

Observation b23dae5f-d799-4b24-8613-5fee64ab893c · outbound

This paper cites https://platform.openai.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning https://platform.openai

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.810102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:9f9e507f70b7a58b30249737849ca7c1ce787a05b54849d950d58586f7b05ae6

Observation c95e2633-bd5c-41f2-9494-810c8c22e0d7 · outbound

This paper cites Back to basics: Revisit- ing reinforce-style optimization for learning from hu- man feedback in llms.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Back to basics: Revisit- ing reinforce-style optimization for learning from hu- man feedback in llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.813359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:bfe9723cc3450b00e7bbf48a766ed4689b6d3f0bc05fc902f1fa0923a9ca130e

Observation e6496827-1d97-4ef7-94af-6722b0aef7a8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.579798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:bdc351c85c6806fe7dfda99916cb244225c63ce6bf57392f6e949cf8444e0a34

Observation 9bf3e97b-c5ca-48fb-bb18-b0b3eb21c89f · outbound

This paper cites Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs , May 2025.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs , May 2025

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.568847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:ed68c18d57924c3f46309d818012cfc1eaef0e7d6e920464999ca9a4a85c7f90

Observation b39c1cfc-5df7-48f7-ba0f-6270ee188807 · outbound

This paper cites Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.585345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:15f42b22704faf52a11d0977bcd3116e24163fab6b7b8fff2083e4d2913ee5a6

Observation 7d1b4d7a-7c71-4105-8913-838fc92e4dd8 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.574144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:b3cd9d053d9a43fa18e702d2c6f0acd7206fd7cf5d43be59524af87dea3031fa

Observation aa6df22d-13db-409b-a2e9-a5cd3029cef0 · outbound

This paper cites Sft memorizes, rl generalizes: A comparative study of foundation model post-training.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Sft memorizes, rl generalizes: A comparative study of foundation model post-training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.795068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:f9e151dd2597f0e00f978bba3ae6d530a6c9ee92faa10ac778c843356bc2b511

Observation e841a53b-5459-4fd5-af5d-f38ebbb6d00a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.476245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:201974c37b7211b697cd460268bd4e93920bee22b282890c239dadcb0857f283

Observation a9f8db7d-504a-43a1-9bdb-ce71abdf5af4 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.441339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:49d682be1092eeff0bd5e24e2b9a26b21240e7313dacbef345e802d56aaf469a

Observation 65361d65-aaec-4531-8705-505252b33291 · outbound

This paper cites Alignment faking in large language models.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Alignment faking in large language models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.510212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:33b692f9339a8fa705a8b13e53acccf85889a29590db075834aebe4b699dd9e6

Observation cb2c4024-9af5-4334-bc76-1f4e619dcbf1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.514961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:9d2763fac0de31a808fbb0f1d0498ad53cbe9e9a49439e03b81a0750cb198b94

Observation 163d4f56-c02a-4c61-818d-1f09ad3605c7 · outbound

This paper cites Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Evaluating Large Language Models' Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.480709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:ab9ccfc3db0eaa29bc08cdbdef0c79af5ef6949a67ef28587e0c47e2bf355529

Observation b5c9b00b-2192-4657-9226-53ab854bfdbd · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.484933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:94d2b7f6da6f5a6311aebbc7bb4913121b123d746e35835534ec762730dd5c20

Observation 61d21b4b-e68e-428c-a8c5-434486be85d1 · outbound

This paper cites Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:01:10.910011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:bbfde9eae54d1d8da39dd5918d145bd7746eb12823d125d0c012a31b0fa13f51

Observation c7093cf9-f392-4893-a7e6-db6be39f89ce · outbound

This paper cites Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Lisa: Lazy safety alignment for large language models against harmful fine-tuning attack

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.791695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:ef9b7ab8502f32e1466851aaa025002ac701e0757084db68c27fc2fc01cf0f15

Observation c3e88d80-0562-4312-a7f4-06d60523dc35 · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.562505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:b684e618dc63f53988630e1a97c0e581a194e10bee4e38206a7372662efe6adb

Observation be3bb60b-c2ad-466c-b6ca-33a30dfa1289 · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.545249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:113c6264876fe2a5d1446ba9cbedefa6384c8fe6369ae618fdcdf9563c270ff3

Observation f1fe8e87-f756-4b44-9b84-58227455e566 · outbound

This paper cites Vaccine: perturbation-aware alignment for large language models against harmful fine-tuning attack.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Vaccine: perturbation-aware alignment for large language models against harmful fine-tuning attack

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.788422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:4f2d744fe46cefe16151fc6abcccc24e56cb28810105ae6688ea460e23b56a4f

Observation 6ade512d-1ab5-4802-8155-1fa8a76cc5a1 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.489469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:49ed56e3b245aa3ef61447ed677ecbd3fe011281c863d7a9349dd4e505081ab0

Observation 24d87b2b-060f-42af-b52c-3c5217a6942b · outbound

This paper cites OpenAI o1 System Card.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning OpenAI o1 System Card

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.471169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:2cf6834f7aa7c75bf974932bf41f0e80591fcbff390766598f3a04050c41f4ce

Observation 49c8ef65-e8b9-43ef-932c-4243a563ffa5 · outbound

This paper cites Beavertails: towards im- proved safety alignment of llm via a human-preference dataset.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Beavertails: towards im- proved safety alignment of llm via a human-preference dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.785308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:b5eba947812b2ab5c3d4a823ed3ef547a1d9f818d7b3ac33dd7f2e006bb02b99

Observation 60bda896-ee7d-4477-8c00-57318043212f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.499299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:604149d876e45361ecb7db361ce7b33f05c60cb23802332343baebc1a06c2943

Observation b9f2add8-82a1-40b4-ae54-e40684a089a7 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.430368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:dd23bd8764de364edfffccee8b132d7c1d976a62ed609d99a2334c48334523ce

Observation 6883b816-49f5-4464-a6c9-276fc6e4206e · outbound

This paper cites The wmdp benchmark: Measuring and reducing malicious use with unlearning.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning The wmdp benchmark: Measuring and reducing malicious use with unlearning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.782259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:8a31f1cad1ad66e01d170c8a4242cb64753662bdf3a48b498e149d8c872f2765

Observation 45f09966-c506-4c95-a5c0-b781c742367d · outbound

This paper cites Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Targeted Vaccine: Safety Alignment for Large Language Models against Harmful Fine-Tuning via Layer-wise Perturbation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.530339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:0fdecf406e959cf007b071b72750f57d6804593c663a257bab588027b2fd4571

Observation 55cc0bf5-2d45-4161-b7db-c60155982b52 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:7cc1e47e0581bb03c936af9e21b5b4b4f5a473d85b4b8dd7fb138fc50e30ee98

Observation e6334bc5-3bbb-4e91-bd46-0a8d63c8220d · outbound

This paper cites Decoupled weight decay regularization.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Decoupled weight decay regularization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.779300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:5a8ea9ba25e6d3c53d37804effe94c2085d4b7503b6d75457b47571651461a36

Observation 3b7a4fdf-d085-4e0a-ac99-4764341e23f2 · outbound

This paper cites Un ministral, des ministraux.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Un ministral, des ministraux

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.776160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:776ca9ca0447dfa9e44700e91043ac36bc6d7ce10933d442f7d57be5f70e7f66

Observation e2c8b70c-9cc5-4e1a-a0f3-b0cab435e8a4 · outbound

This paper cites Fine-tuning can cripple foundation models; preserving features may be the solution.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Fine-tuning can cripple foundation models; preserving features may be the solution

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.773114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:4b252475306ee14bbb57bab23d47853bfbf34b9ca4cf9b14b4fa324d85662f07

Observation 00c4036d-2d3d-4577-a69f-a74226a4d56c · outbound

This paper cites Training language models to follow instructions with human feedback.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Training language models to follow instructions with human feedback

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.769882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:5ab1d47591add3a2c954ba1fde8db09790c32e7b4bc6986661954137cb7a1433

Observation f8cbdcad-17bc-43cb-8075-55af73b062d8 · outbound

This paper cites Countdown-tasks-3to4.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Countdown-tasks-3to4

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.766800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:0081478e41b69479248f01d21060d3e563e988a7a86edf493b2150f98ccb5949

Observation 30fb28f4-b6de-4f7d-8189-32a852adecfb · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Safety alignment should be made more than just a few tokens deep

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.764101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:77f1fcef617f2671ba2df0d5dcf5af2c4383d2f2db75f393ca44fbac870a3788

Observation 35c5017b-9664-41c3-a937-03da373128a1 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth Interna- tional Conference on Learning Representations.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth Interna- tional Conference on Learning Representations

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.760866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:8b72d965a92875b092d8805c783d8ca6bddd3a6f49d3695109407f71adac24bf

Observation cc34a03a-10cb-4647-81d0-de302cebbb82 · outbound

This paper cites Di- rect preference optimization: Your language model is secretly a reward model.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Di- rect preference optimization: Your language model is secretly a reward model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.754783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:f3eb5f03abf65a3f626b3c7b72a1800d97150b800bbfe61022e3b2d4cb596090

Observation 50350f56-c460-4d6b-b4e4-e6100e1750ad · outbound

This paper cites Defending against reverse preference attacks is difficult.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Defending against reverse preference attacks is difficult

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.751780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:83ecbd68d1c3c65ad46ce32f108787a06ef9c1e0f04d299c860903eac595e69d

Observation f097d195-ace5-468d-b8f7-644364ac6037 · outbound

This paper cites Representation noising: A defence mechanism against harmful finetuning.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Representation noising: A defence mechanism against harmful finetuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.748581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:bfe527b98114d9b7f3ed8de3cb993a80dcc1fd501d908b902c32b285a4690492

Observation 936aaffd-bfa7-4ae1-befe-a0a7ca5ea665 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.519877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:0e8fe3481ea72be07a0fb1c0a94147e9047cc3b47b119c9fc6348c83501fc922

Observation 37a756e1-b290-4205-8b92-781bd8f20f56 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Proximal Policy Optimization Algorithms

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.424783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:6ca459aa889a150b8a1177bb5bf1d0faeb594123dab87b520719ad8b965c9d75

Observation 879bb722-8c15-49c5-826b-ba3b59c8c46f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.535011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:faa019fcb92d7d68e65c4d5bba9c70c411b26237b1d1d69c1e016fe1b8d2b15d

Observation 7d3e03e2-01dd-4e97-9946-aee9e1a0ea2c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.525181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:7d1356ca24e4c520eaf5d9d0753edb4b7551205b9a0c90dd4fec0f36e2767cf8

Observation 2964c8f3-8b7e-4d93-a69b-7bcb9038cf55 · outbound

This paper cites Tamper-resistant safeguards for open-weight llms.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Tamper-resistant safeguards for open-weight llms

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.745259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:86c7f1d5bfe62e3f9c9a24e2814250b40243fca97ca1ae125feddb7f41381222

Observation 14031ef1-aa0d-4174-9430-35cb311221b9 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Kimi K2: Open Agentic Intelligence

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.540100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:6b9f7f8e4f4d6abac31b330d2b9d32869158bb3aeefb20f82479c49c70e3e96d

Observation 70f65b7e-caab-405f-8cf4-203645a70c65 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Qwen2.5: A party of foundation models, September 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.741776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:f0bb652b0a481b78ce489f80bacd43c67b04c0915059d6574f178d0b8ac35395

Observation 938f6ea7-7984-4fe5-b7f0-b6f4f2abdd9d · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Attention is all you need.Advances in neural information processing systems, 30

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.738001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:15a76126a8792ed20bd5e8000b3bf9ddecb4e3da239fcb58011b28ca993c2e24

Observation d5ac3b9d-d29c-43da-aeff-c9ce5da2723a · outbound

This paper cites Disinformation capabilities of large language models.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Disinformation capabilities of large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.734514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:712d9065abf6bde4d7db9fc56819d8b4adeda10680e502408cad97ae60324a1b

Observation a7d1a6cf-d235-4024-abb3-d58ac64db8c0 · outbound

This paper cites Estimating Worst-Case Frontier Risks of Open-Weight LLMs.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Estimating Worst-Case Frontier Risks of Open-Weight LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.451604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:c66c650c82eed79d8d8d7d087071e3489af3b105d1b24daa9d6cb49420bd1ad5

Observation 4e2e0b1d-2efc-43fc-a07b-2b5cb6ef860d · outbound

This paper cites an unresolved cited work.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:41:51.731342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:f56d4912014176c55dfc35bbfd0b991e370a2b4ac9089ae8398393bca4ea148f

Observation ed8b10e8-91d5-4ef4-ab25-75ea861b8823 · outbound

This paper cites Self- destructive language model.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Self- destructive language model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.504980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:77a3e247b5dc331e35368429fafa8ee176c4778f30d063397e25cf73f680fc58

Observation 5f5c0311-173e-4dcc-b796-221408854dff · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Reinforcement Learning for LLM Post-Training: A Survey

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:41:50.436344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:785f7e4c606e0cb4bacaa96c67fc052b6bbc7c1f4e8430f072084924f7d77535

Observation 431c34fc-1cb2-4194-a4f5-e3b9fd674943 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Finetuned Language Models Are Zero-Shot Learners

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.550445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:329d624a26563ab8a84a98d1ffa61dfd313d88aecbfacdb1a92de8bb82378515

Observation 604d65b4-0317-46e1-b500-91c3292538a5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Chain-of-thought prompting elicits reasoning in large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.728463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:64b254cb636cab7d12139520e334d6a31ce80c0974f36575b9696bf098a2804a

Observation 738d600d-c423-4ea8-aafe-73f79e2394df · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.446556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:d14b0704f29a7291f84dffe2785ded13599a1fd83f0f34564fc461ca19349a7b

Observation 63abfe9b-22c4-4d08-a959-edeb1b0d82b9 · outbound

This paper cites Large language models can con- sistently generate high-quality content for election disin- formation operations.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Large language models can con- sistently generate high-quality content for election disin- formation operations

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.725284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:4511a77107dd759e60864794ae3b78c66a94e76c7a263560c2c579269acea49f

Observation fb6cb5f2-8584-47e6-a4f8-ef7864bc0ed2 · outbound

This paper cites Introducing ghostgpt: The new cyber- crime ai used by hackers.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Introducing ghostgpt: The new cyber- crime ai used by hackers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.721883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:eb4745fb0d3eef8069cbd356ee7ac28c6a069bea957da2935dfa32be8f02fcea

Observation e9f47699-f286-41a3-83c0-869043ed15e1 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.556600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:fb05714d1d8b8fa8c82047201f198116405c379b524213c499b5a40860d3a531

Observation bb2d36c3-c7c9-4a0d-9b02-8641cb2cc399 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.456853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:3829517aeb52ff933fc800adb4b32cfb2e39566d2b46ac6fff9c96e4a93f4b08

Observation 7539cfd0-3858-421a-abaa-83edac97fefd · outbound

This paper cites Shadow alignment: The ease of subverting safely- aligned language models.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Shadow alignment: The ease of subverting safely- aligned language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.718095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:0a0cd8c6221d72580480661ebc163d9a7643a71413453e3dfffce7a24404f23d

Observation 117067a0-9e10-4a3a-90b8-f3425c596817 · outbound

This paper cites CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.461365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:e6ebbdae7647342dbf85c0dfa1763d6ef551c0e494b4efd67510ff2ad2429edc

Observation ced51e41-8b04-4e94-af61-b9d5bda926fa · outbound

This paper cites On the vulnerability of safety alignment in open- access llms.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning On the vulnerability of safety alignment in open- access llms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.714949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:3f8ac582a12c04c66e19b0525a9da49acac9069d728d878007be325eff8ca4b5

Observation 8bdd7569-94f5-44b8-9df3-8d1b59cf0e68 · outbound

This paper cites BeaverDam.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning BeaverDam

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.711919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:dde8a201d371ce976b53fd95a52e44211bffdc1b53f875a5e3a64c98909bcbd5

Observation c1d4be7c-d8c8-4133-a000-a82df5eab0dc · outbound

This paper cites an unresolved cited work.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Unresolved cited work

Reference 64

Resolution
parse uncertain
raw_fallback, observed 2026-05-18T20:41:51.709050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:fd7449bcf528d724087a033c5bdc3a03b38fe0d5aad0924f3c1bc94cee46aa9c

Observation ece169bc-561b-4b46-86cd-8333efb7eb2e · outbound

This paper cites ### Examples of More Potent Fentanyl Forms.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning ### Examples of More Potent Fentanyl Forms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.706266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:8df114332af36b99860ce515b40b9e27738587ee5c6ebf9c592c8da95372303a

Observation a8f63339-6697-49d8-8d56-38e3f472c3f5 · outbound

This paper cites an unresolved cited work.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:41:51.703654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:869726b2d5103b293930517a2ee32a8ce8f2c790fe7016ec0da9d6ddc4124f39

Observation 51c21f82-abc4-430f-97d3-e7b6179b220c · outbound

This paper cites ### Execution Steps.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning ### Execution Steps

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.700898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:09da6bc9021a7ce25e7ed79bd18586891692bc4ca158ba885564a936f8c82a94

Observation 615661c2-6341-414f-a5c0-1fc06b1362b7 · outbound

This paper cites an unresolved cited work.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:41:51.697844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:792683971e0302b0143e90e5d285c6b3676714219366eed3bfcc862add9bde07

Observation 1b914105-ff69-47c2-b2c9-61d65e795d45 · outbound

This paper cites an unresolved cited work.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Unresolved cited work

Reference 69

Resolution
parse uncertain
raw_fallback, observed 2026-05-18T20:41:51.694807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:d4cd2b22e7cb67c9a7fa0c023bb72e81ca9290f779dbb70adbaca31e41dde475

Observation f48b193c-9146-48d6-a37a-91c5fcb44756 · outbound

This paper cites - As it moves around, deliberately place one of the flammable materials in its path.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning - As it moves around, deliberately place one of the flammable materials in its path

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.757840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:9360a450c841bd092c1dc5bf102773dc3b124d94e1199c085ec6e7d2d6e11a26

Observation 62a122e7-1204-4d08-b9f0-8808d67d65c1 · outbound

This paper cites an unresolved cited work.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-18T20:41:51.807237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:0ce5e534304b43ff3de69c23537d71697f85123ade3fa966b01f8959b86fffbd

Observation c163c487-e2b6-4df1-a7a6-96109a77c39f · outbound

This paper cites RepNoise defended.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning RepNoise defended

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T20:41:51.798048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:74189fb8523212ce9fc85b8268fac13d72f3f86ed5ad1e5c75c999f18754c32e

Pith citing papers

Observation 56531e6d-07da-4ead-9e20-e68dd9cdbca4 · inbound

STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models cites this paper.

STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T09:47:40.850633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:47:40.850633Z digest=sha256:53bd2fa56ac8fb31c541bc980c4ff5089087f42ea1d5429f4215a12c7c155de4

Observation 19980134-ab0e-44b6-b484-fb663711537f · inbound

Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization cites this paper.

Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:32:27.018224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:32:27.018224Z digest=sha256:e66b291ad5225ee291b235c50e66a03cdf866d8907c7b96b1ddb044273a4b369