Pith. sign in

Paper Citation Record · LEDGER

Crafting Reversible SFT Behaviors in Large Language Models

As of 6 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2605.06632.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06632 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:16:16.104176Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact22
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 212cdefd-ab25-4d43-a896-082f513fb478 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

Crafting Reversible SFT Behaviors in Large Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.118398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:37248e9ec069cde04cb1585dbc2480764d8dc72daf879f1dcaeff78edb5df793

Observation 90dba320-60ff-4b57-afb0-efacd5f239e4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Crafting Reversible SFT Behaviors in Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.404236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:d94763cdb752d27af2556d9c7e8362ebe9c8953853b0d79ff8d516d963774e81

Observation 4c6a3390-5255-4fca-8d7a-37e77c6fe711 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Crafting Reversible SFT Behaviors in Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.371175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:a1db5705990163aa09818bf71c7f2decd783ca8185c18919e2a58eae08eed62c

Observation cb8e3a34-17c1-4e03-95ac-82b2e70782a3 · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021.

Crafting Reversible SFT Behaviors in Large Language Models Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.177787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:0fb6147f8086d05ab947facfb895cd5c95282e691999fdbd971e8f8c19722640

Observation a65cce11-f0b7-4f7e-ba3d-98bab8616417 · outbound

This paper cites Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning.

Crafting Reversible SFT Behaviors in Large Language Models Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.395858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:257c2d5e2282546a6c1f843198a712d9762052fa0ce495337e95eae0f109a34a

Observation 5ccce7eb-640a-45f7-9be5-5d5d8433e00e · outbound

This paper cites A mathematical framework for transformer circuits.Transformer Circuits Thread, 1(1):12.

Crafting Reversible SFT Behaviors in Large Language Models A mathematical framework for transformer circuits.Transformer Circuits Thread, 1(1):12

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.169146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:e975e325347bf6574982ec88bcac49f423def16409e983dd0d081908445be0cc

Observation 18288c93-ae68-4879-a13e-c728050440cc · outbound

This paper cites Zoom in: An introduction to circuits.Distill, 5(3):e00024–001.

Crafting Reversible SFT Behaviors in Large Language Models Zoom in: An introduction to circuits.Distill, 5(3):e00024–001

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.173112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:d63f984cd42eb001611fcd04466fe8dbaabaf8293644a04a076c80328b2f0a4b

Observation e740ec2a-57c6-4e27-903c-8f3cc2052f42 · outbound

This paper cites Superficial safety alignment hypothesis.

Crafting Reversible SFT Behaviors in Large Language Models Superficial safety alignment hypothesis

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.425920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:f9f9451cbc16e7ea01a8975d045bdc57184d3f68030485e6ecd925160f5bb70f

Observation 1fb1d1fe-3c8d-48e7-927d-832f92214025 · outbound

This paper cites A Layer-wise Analysis of Supervised Fine-Tuning.

Crafting Reversible SFT Behaviors in Large Language Models A Layer-wise Analysis of Supervised Fine-Tuning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.289378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:2688d9ec33d9128f716fc6f2e8147adb76a0c4e91a62c2c212a21a8313143183

Observation e998f420-3d1f-48a3-8d52-0df82a4aa6f4 · outbound

This paper cites Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns.

Crafting Reversible SFT Behaviors in Large Language Models Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.296496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:d00a7d8e631ae3abc2b7c681ef381c6a3e57dffc4cdba71faf52e1cddd2f2a3f

Observation 1991c15c-c5da-44ed-bfc2-23ccbf076010 · outbound

This paper cites Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting.

Crafting Reversible SFT Behaviors in Large Language Models Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.320848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:0c868f522f20afed96d59af22b1977916fa3f0e3035c8eb4f11eeef0ca7fff47

Observation 6ea2c7dc-3824-4aa3-b692-d98d60b5a604 · outbound

This paper cites Talking to yourself: Defying forgetting in large language models.

Crafting Reversible SFT Behaviors in Large Language Models Talking to yourself: Defying forgetting in large language models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.422105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:ddec934e8b41f63d586ec00ffa13f82081455db50f696bf8c0bc995c74997aaf

Observation 4892f287-79a0-4f44-9379-3d1721e0090e · outbound

This paper cites UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function.

Crafting Reversible SFT Behaviors in Large Language Models UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.387461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:51f5343cc00b7df4b98a422e106738a23fdf6098f756568f06b166f4ac81dba1

Observation 23dbe324-e913-475f-a7a1-128b78ff4bb7 · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.Advances in Neural Information Processing Systems, 36:16318–16352.

Crafting Reversible SFT Behaviors in Large Language Models Towards automated circuit discovery for mechanistic interpretability.Advances in Neural Information Processing Systems, 36:16318–16352

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.181721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:281c8809d1714c99f9e230b6c2a7dec5f9894102642bf88db8f4e3d49c55e471

Observation aa4ca53f-db40-4674-8057-fff431629457 · outbound

This paper cites Attribution patching outperforms automated circuit discovery.

Crafting Reversible SFT Behaviors in Large Language Models Attribution patching outperforms automated circuit discovery

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.185413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:52cfed2dcb052247967c4be7618d089f0aa5997fe02170d885c1adcaff110830

Observation 659057b9-81c7-4631-a39f-844ffc75e535 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Crafting Reversible SFT Behaviors in Large Language Models Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:15:10.863432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:0586a5c9e73ab939ae5f836095db2c0e726f00d37244a3c60a96fbdd5955201d

Observation e5455b24-f7ab-4342-919f-4ee47cc19831 · outbound

This paper cites Safeseek: Universal attribution of safety circuits in language models.

Crafting Reversible SFT Behaviors in Large Language Models Safeseek: Universal attribution of safety circuits in language models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.356055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:0675817757354f41f281311309a65c458703432c3eb18eff1f41c642ff571e43

Observation 06d07e09-72d9-431d-8302-f06b2f15887d · outbound

This paper cites arXiv preprint arXiv:2511.13653 , year =.

Crafting Reversible SFT Behaviors in Large Language Models arXiv preprint arXiv:2511.13653 , year =

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.439104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:b8c2c9f80b89bced27ec5bd750a176d57311076358fc42b27aed31d3efc28799

Observation f84eb357-4982-4076-bcb0-657654d712f1 · outbound

This paper cites Toy Models of Superposition.

Crafting Reversible SFT Behaviors in Large Language Models Toy Models of Superposition

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:44:43.591728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:8060da168c3316b0f79d1cbc6860bf85fa349a499e1e7121259d20380657abd9

Observation 1d593131-372a-430c-8133-b1ab58d834b6 · outbound

This paper cites Editing Models with Task Arithmetic.

Crafting Reversible SFT Behaviors in Large Language Models Editing Models with Task Arithmetic

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:09:13.515542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:f8d3018e200f2044936c4a8efddb90260dce356fe933405b8dbe43fe7aa4d6f4

Observation 8a7e203f-707d-4472-96a2-5a111773921a · outbound

This paper cites Task arithmetic in the tangent space: Improved editing of pre-trained models.Advances in Neural Information Processing Systems, 36:66727–66754.

Crafting Reversible SFT Behaviors in Large Language Models Task arithmetic in the tangent space: Improved editing of pre-trained models.Advances in Neural Information Processing Systems, 36:66727–66754

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.189441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:42cb8a7f8dd5af11ebe21fa58ba011957f53e940ad2ec186d1d99e497bab2e1c

Observation 1f383852-6c70-4a86-bd40-829a667ea313 · outbound

This paper cites Causal scrubbing: A method for rigorously testing interpretability hy- potheses.

Crafting Reversible SFT Behaviors in Large Language Models Causal scrubbing: A method for rigorously testing interpretability hy- potheses

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.145858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:249f38c15ff3a0836870bc1b32183a5fae8d1ff2a9ffb7cc9519c0dcd12fd04c

Observation 38972c5f-4ead-450b-993f-caf458c50ca3 · outbound

This paper cites Badnets: Evaluating backdooring attacks on deep neural networks.Ieee Access, 7:47230–47244.

Crafting Reversible SFT Behaviors in Large Language Models Badnets: Evaluating backdooring attacks on deep neural networks.Ieee Access, 7:47230–47244

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.149313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:4e27d100d44bf012d774d54532bc5d707914bada558aa0b7f780242f8671f78a

Observation 484f6eb9-50e4-49e1-8cfe-6b2a792ea1f1 · outbound

This paper cites Depth charge: Jailbreak large language models from deep safety attention heads.

Crafting Reversible SFT Behaviors in Large Language Models Depth charge: Jailbreak large language models from deep safety attention heads

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.341152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:b8dee53488dcc43d75c91f9411359493244f9fbcc8e05dd3b6d4e45b1df7c0cc

Observation 0ce29a2c-ee2b-4fa1-8261-2d2ea214c04e · outbound

This paper cites Piggyback: Adapting a single network to multiple tasks by learning to mask weights.

Crafting Reversible SFT Behaviors in Large Language Models Piggyback: Adapting a single network to multiple tasks by learning to mask weights

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.160180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:bab7dc26dcfc5bdcaddeeeab978a48d9fe428eb280c0ce0e3389f0554c95e625

Observation 44b45d82-b79a-495b-8e22-8678d711f958 · outbound

This paper cites Learning Sparse Neural Networks through $L_0$ Regularization.

Crafting Reversible SFT Behaviors in Large Language Models Learning Sparse Neural Networks through $L_0$ Regularization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.312678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:51e106e609af095421ea95df4ba40b9777512ea193597e84ed538db15bff54ce

Observation 1e97eeff-f62a-431b-be18-cbb34a5abf0f · outbound

This paper cites Low-complexity probing via finding subnet- works.

Crafting Reversible SFT Behaviors in Large Language Models Low-complexity probing via finding subnet- works

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.130187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:c5fd703e5d3f4384494eb4c464499a2fd2a153d1c056475b27fd8351d64363d8

Observation ee53de27-c53f-4149-bd56-4220e7a66f91 · outbound

This paper cites Hashimoto.

Crafting Reversible SFT Behaviors in Large Language Models Hashimoto

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.133590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:d4daf70dc0da500f11309e0a93d780442a55b3e2a2927d1945f2d539154734e6

Observation 271da2ca-dc8c-4368-b01b-05b9660aa03a · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165.

Crafting Reversible SFT Behaviors in Large Language Models Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.137673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:7a99484386f0881dfbb15ffbdad09e6cfd9a398079239168f0bc71a3518cc4a2

Observation acdb678b-5f67-4b9a-a839-61082dae3b10 · outbound

This paper cites Shakespearean and modern english conversa- tional dataset.

Crafting Reversible SFT Behaviors in Large Language Models Shakespearean and modern english conversa- tional dataset

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.141703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:f33b63e77f223e7e677644bfa8c8bc350a975fb6efcb26f2e593dfce18fc6933

Observation 06e4be68-cbbb-466f-a7cc-ea4839982861 · outbound

This paper cites Qwen3 Technical Report.

Crafting Reversible SFT Behaviors in Large Language Models Qwen3 Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.348643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:b551fc8e1f3503d2a734065144dd6d46f380289b6a1cb1be0e5d7b3d63a80c3e

Observation d20c2cd6-21fe-4d8b-8021-f62a5a12ac57 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Crafting Reversible SFT Behaviors in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.333193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:f36154034f8400d0f4ad4adaa07e1b93ca96ceccc12dcb56a8f451cd17f52618

Observation 012ebb21-4f98-49aa-a3dc-693b21475650 · outbound

This paper cites Mistral 7B.

Crafting Reversible SFT Behaviors in Large Language Models Mistral 7B

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.306114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:2a09ec502a4d449cf70abb58fe8f587550503be1ce0ed97acf6762213b411efa

Observation a1e189f5-fda2-46ac-b562-c4befa0e316a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623.

Crafting Reversible SFT Behaviors in Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.153708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:ca462fe31004982773ac5f704d493712d76f97f41c6f12f97fbc340ce015a2b0

Observation fcb1a335-6b1f-4a98-a3de-0c06f509f439 · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in neural information processing systems, 37:8093–8131.

Crafting Reversible SFT Behaviors in Large Language Models Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in neural information processing systems, 37:8093–8131

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.164843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:7520485aafcd26473a5aeb374a4c26cb1056b9d7a919e367866262c1028bf038

Observation 83842c3e-3f6a-48a8-ba8d-9464410eb82b · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Crafting Reversible SFT Behaviors in Large Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.429034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:c22e66ca3e41a233b3fb6ec42ac8b0060a460760e381991672511d1cbdf25a5c

Observation 02c1b365-b199-4629-aeb3-99c5f8015018 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Crafting Reversible SFT Behaviors in Large Language Models Measuring Massive Multitask Language Understanding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.414568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:d24352286460bd6b9a15bbb727a3148b0a092f73b255921dd41775de6467108e

Observation cf97c592-83b0-40ca-97ef-7203e2ff7296 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800.

Crafting Reversible SFT Behaviors in Large Language Models Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.122206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:82bc64fff665c877e860b5cb4409c3429b07ace184c20baf1f5f04c1c6f1972f

Observation b30aa9f4-834b-4d6d-a90c-b7a56785b16d · outbound

This paper cites Modeling by shortest data description.Automatica, 14(5):465–471.

Crafting Reversible SFT Behaviors in Large Language Models Modeling by shortest data description.Automatica, 14(5):465–471

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:32:26.126947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:0706a31b1f61468692dbbb93e9e75bc88314605f76b412b8a78b097d245a2416

Observation 80d95ea0-c554-450a-b223-8f4359fd07f8 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Crafting Reversible SFT Behaviors in Large Language Models Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:21:07.383433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T12:16:16.104176Z digest=sha256:b80b3278fb0713ce0543bf39d85c088ec584ab0753c4998b9d3e9a5a864b359d

Pith citing papers

No inbound Pith citation observations are available.