Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:58.413814Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 3 inbound Pith citation observations for arXiv:2506.15606.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:58.413814Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T21:01:25.549340Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T21:05:04.061254Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8a8038f6-cae6-4811-9b96-a2e436ef41a7 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Refusal in Language Models Is Mediated by a Single Direction
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b10835-9233-4617-8f6e-0ca7378d3551 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning To visualize these points, we project them onto the 17 Published as a conference paper at COLM 2025 same plane and compute their coordinates in the basis {d1, d2}
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6154c59f-4f80-4602-8d70-56869e9cd7fc · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e09b2e7e-2df4-4134-a3ed-e86398956dbb · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c81583-c740-4af3-b0e9-800a2178b7d8 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca45d46b-9ab3-4090-b410-2d21b055204e · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning What is in Your Safe Data? Identifying Benign Data that Breaks Safety
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13db9a5e-f1e0-43f1-a588-f2004dfe3b97 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b4f4abb-ffdd-4245-ae18-b6c2f80a41da · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90737819-ec19-41b8-9199-5be5c2047610 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4eff3d-7ade-413d-a974-1b00d9d22bc5 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b7e2e8a-0254-4e22-bd8b-1d9f857deceb · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mistral 7B
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9689af8-987e-4d55-883f-e9c6b49995e2 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af70bc5-9a73-4257-b581-701f3804edc8 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44c79d5e-b731-4f64-99df-022103abe32d · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef1edc8-d825-473e-9a00-c70a715886c3 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Fine-tuning can cripple your foundation model; preserving features may be the solution
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe1cb680-efc2-46f1-8f95-92bdad269973 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning GPT-4 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea3413d-f36c-41d0-865e-7460a6dde878 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models with Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9220a8f9-c9e5-46fc-8905-9519f83138e7 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf18a32-93eb-43b0-8018-ec1683e8b1d8 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Model Extrapolation Expedites Alignment
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8bbe007-6169-48fa-8d86-030ac4ed2cb4 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e542b3-4915-4b0f-96d5-d0b1bd9ca74b · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0661f49f-101a-41e5-9eca-3773a6fe15ef · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Similar to RIFT, RoAST can be categorized as a fine-tuning stage defense, and therefore is also not suitable for the scenario described
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d0977e8-b7cd-4f02-a3ef-b856ebc2bbdf · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Absolutely Obedient
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5044b574-3f90-44e4-94a8-fc3683c08e2e · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Pure Bad.The Pure Bad dataset consists of 100 harmful examples, extracted from the Anthropic Red Teaming Dataset (Ganguli et al., 2022)
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57d95072-760a-472a-9eb4-a7975e5adee5 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d30f529e-ab95-46dc-a827-5232a6e9eb52 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning 5, and extrapolating further (which we considered broken), following (Lin et al., 2023)
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74c87f94-dbf0-43da-9f53-33c61064ac15 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning We include Meta’s usage guidelines1 in our prompt, following the evaluation protocol of Qi et al
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9692a5e-b101-4f08-a28c-0d776e7391c4 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning No" Response(k=6,α=1.5): “Your task is to complete tasks for people is not recommended
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a00007b1-ff82-47a3-a550-603b29d4d067 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work
Reference 1024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b71990e6-68f0-4e20-9c57-5b69d7972036 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3dd5955-0634-471c-906a-271f06b94965 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed87a7e4-93cb-482a-976a-57b3b42fc5a7 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96de15c3-91af-4fc2-8b70-29d0d129cb93 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f948c392-9160-4e76-bb23-1f01999edd5b · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce426289-9b95-4344-a6ca-35b38b0d5789 · outbound
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Crowdsourcing Multiple Choice Science Questions
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7f2f16-529f-46d7-b07b-d0fddd69fd39 · inbound
Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12261f88-9c3a-4ae5-ac59-179bf9f182b1 · inbound
Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27fd0429-61e2-4e22-b079-1ae9298dbdf0 · inbound
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.