Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T12:16:16.104176Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2605.06632.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T12:16:16.104176Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 212cdefd-ab25-4d43-a896-082f513fb478 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90dba320-60ff-4b57-afb0-efacd5f239e4 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c6a3390-5255-4fca-8d7a-37e77c6fe711 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb8e3a34-17c1-4e03-95ac-82b2e70782a3 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a65cce11-f0b7-4f7e-ba3d-98bab8616417 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ccce7eb-640a-45f7-9be5-5d5d8433e00e · outbound
Crafting Reversible SFT Behaviors in Large Language Models A mathematical framework for transformer circuits.Transformer Circuits Thread, 1(1):12
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 18288c93-ae68-4879-a13e-c728050440cc · outbound
Crafting Reversible SFT Behaviors in Large Language Models Zoom in: An introduction to circuits.Distill, 5(3):e00024–001
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e740ec2a-57c6-4e27-903c-8f3cc2052f42 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Superficial safety alignment hypothesis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1fb1d1fe-3c8d-48e7-927d-832f92214025 · outbound
Crafting Reversible SFT Behaviors in Large Language Models A Layer-wise Analysis of Supervised Fine-Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e998f420-3d1f-48a3-8d52-0df82a4aa6f4 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1991c15c-c5da-44ed-bfc2-23ccbf076010 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Improved Supervised Fine-Tuning for Large Language Models to Mitigate Catastrophic Forgetting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6ea2c7dc-3824-4aa3-b692-d98d60b5a604 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Talking to yourself: Defying forgetting in large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4892f287-79a0-4f44-9379-3d1721e0090e · outbound
Crafting Reversible SFT Behaviors in Large Language Models UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23dbe324-e913-475f-a7a1-128b78ff4bb7 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Towards automated circuit discovery for mechanistic interpretability.Advances in Neural Information Processing Systems, 36:16318–16352
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa4ca53f-db40-4674-8057-fff431629457 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Attribution patching outperforms automated circuit discovery
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 659057b9-81c7-4631-a39f-844ffc75e535 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5455b24-f7ab-4342-919f-4ee47cc19831 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Safeseek: Universal attribution of safety circuits in language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06d07e09-72d9-431d-8302-f06b2f15887d · outbound
Crafting Reversible SFT Behaviors in Large Language Models arXiv preprint arXiv:2511.13653 , year =
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f84eb357-4982-4076-bcb0-657654d712f1 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Toy Models of Superposition
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d593131-372a-430c-8133-b1ab58d834b6 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Editing Models with Task Arithmetic
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a7e203f-707d-4472-96a2-5a111773921a · outbound
Crafting Reversible SFT Behaviors in Large Language Models Task arithmetic in the tangent space: Improved editing of pre-trained models.Advances in Neural Information Processing Systems, 36:66727–66754
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1f383852-6c70-4a86-bd40-829a667ea313 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Causal scrubbing: A method for rigorously testing interpretability hy- potheses
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 38972c5f-4ead-450b-993f-caf458c50ca3 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Badnets: Evaluating backdooring attacks on deep neural networks.Ieee Access, 7:47230–47244
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 484f6eb9-50e4-49e1-8cfe-6b2a792ea1f1 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Depth charge: Jailbreak large language models from deep safety attention heads
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ce29a2c-ee2b-4fa1-8261-2d2ea214c04e · outbound
Crafting Reversible SFT Behaviors in Large Language Models Piggyback: Adapting a single network to multiple tasks by learning to mask weights
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44b45d82-b79a-495b-8e22-8678d711f958 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Learning Sparse Neural Networks through $L_0$ Regularization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1e97eeff-f62a-431b-be18-cbb34a5abf0f · outbound
Crafting Reversible SFT Behaviors in Large Language Models Low-complexity probing via finding subnet- works
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee53de27-c53f-4149-bd56-4220e7a66f91 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Hashimoto
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 271da2ca-dc8c-4368-b01b-05b9660aa03a · outbound
Crafting Reversible SFT Behaviors in Large Language Models Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation acdb678b-5f67-4b9a-a839-61082dae3b10 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Shakespearean and modern english conversa- tional dataset
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 06e4be68-cbbb-466f-a7cc-ea4839982861 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Qwen3 Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d20c2cd6-21fe-4d8b-8021-f62a5a12ac57 · outbound
Crafting Reversible SFT Behaviors in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 012ebb21-4f98-49aa-a3dc-693b21475650 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Mistral 7B
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1e189f5-fda2-46ac-b562-c4befa0e316a · outbound
Crafting Reversible SFT Behaviors in Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fcb1a335-6b1f-4a98-a3de-0c06f509f439 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in neural information processing systems, 37:8093–8131
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83842c3e-3f6a-48a8-ba8d-9464410eb82b · outbound
Crafting Reversible SFT Behaviors in Large Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02c1b365-b199-4629-aeb3-99c5f8015018 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Measuring Massive Multitask Language Understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cf97c592-83b0-40ca-97ef-7203e2ff7296 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th annual meeting of the association for computational linguistics, pages 4791–4800
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b30aa9f4-834b-4d6d-a90c-b7a56785b16d · outbound
Crafting Reversible SFT Behaviors in Large Language Models Modeling by shortest data description.Automatica, 14(5):465–471
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 80d95ea0-c554-450a-b223-8f4359fd07f8 · outbound
Crafting Reversible SFT Behaviors in Large Language Models Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.