Pith. sign in

Paper Citation Record · LEDGER

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 3 inbound Pith citation observations for arXiv:2506.15606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15606 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:58.413814Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:01:25.549340Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.061254Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a8038f6-cae6-4811-9b96-a2e436ef41a7 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Refusal in Language Models Is Mediated by a Single Direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.404563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.404563Z digest=sha256:205068030b07a8929d88833e979a83d58d400be3fc61dd51d48e1af3cca4e3b4

Observation 38b10835-9233-4617-8f6e-0ca7378d3551 · outbound

This paper cites To visualize these points, we project them onto the 17 Published as a conference paper at COLM 2025 same plane and compute their coordinates in the basis {d1, d2}.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning To visualize these points, we project them onto the 17 Published as a conference paper at COLM 2025 same plane and compute their coordinates in the basis {d1, d2}

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.628892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:58.196559Z digest=sha256:a378411fd738d4abab07330a481d9fe53db7504eb76ddf31f853bc79b07a054d

Observation 6154c59f-4f80-4602-8d70-56869e9cd7fc · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:59:58.836084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:57.924221Z digest=sha256:f5206014f4ff8cde5001ed9b0266f62eb856efdd4c2570bd3911151e2c346216

Observation e09b2e7e-2df4-4134-a3ed-e86398956dbb · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.625447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.625447Z digest=sha256:dc25f57c96891f497a6b47baad1d2dd300f063c1ba6465dbaafe4382f2930c59

Observation 16c81583-c740-4af3-b0e9-800a2178b7d8 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.772459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.772459Z digest=sha256:a6acf59d1f77492fc14c64506226f2225caba7023899a64e329229dba6ebd59b

Observation ca45d46b-9ab3-4090-b410-2d21b055204e · outbound

This paper cites What is in Your Safe Data? Identifying Benign Data that Breaks Safety.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning What is in Your Safe Data? Identifying Benign Data that Breaks Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.836898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.836898Z digest=sha256:4e03f6567044ed0a980bad0b4873cb7e9e010a9351d5533ed1f144031350d662

Observation 13db9a5e-f1e0-43f1-a588-f2004dfe3b97 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.892653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.892653Z digest=sha256:a64fdf3cd8fb28d64298b67372aff51ecfc2d5cc748b94529c4b6fc48a959bd4

Observation 6b4f4abb-ffdd-4245-ae18-b6c2f80a41da · outbound

This paper cites TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning TextGrad: Advancing Robustness Evaluation in NLP by Gradient-Driven Optimization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:59:59.643808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:55.939806Z digest=sha256:8cd77e13a96ddae53b21866a58d708fb5e80611be6e8da1c2639507f87fde8da

Observation 90737819-ec19-41b8-9199-5be5c2047610 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.005326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.005326Z digest=sha256:ab6b180f27e4a06582ef1881408acf4126f926df294ef645c4ad91148485d176

Observation fa4eff3d-7ade-413d-a974-1b00d9d22bc5 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.106685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.106685Z digest=sha256:9d14f3acf8380c47cfc4c8a85bdba4d7a89376df318097e548e806af28e78534

Observation 7b7e2e8a-0254-4e22-bd8b-1d9f857deceb · outbound

This paper cites Mistral 7B.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.214328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.214328Z digest=sha256:6c69a07e9d2e25906d99d01c534066ca527966fb7c017d2355b8cd4e79bd9cf6

Observation f9689af8-987e-4d55-883f-e9c6b49995e2 · outbound

This paper cites The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.327320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.327320Z digest=sha256:5c9b63ab13f9f3b18d4d68da6b150237fe447d1f17d5d1ec8b0bbbd4b3f087f7

Observation 4af70bc5-9a73-4257-b581-701f3804edc8 · outbound

This paper cites Robustifying Safety-Aligned Large Language Models through Clean Data Curation.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.430016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.430016Z digest=sha256:4134340162d125e38a071e2a5e3dd3e7e39e3f706346a995ff979a75269c4af9

Observation 44c79d5e-b731-4f64-99df-022103abe32d · outbound

This paper cites Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Use of LLMs for Illicit Purposes: Threats, Prevention Measures, and Vulnerabilities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.464742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.464742Z digest=sha256:6fc35d2aa19929178f4474a80168f967feabb7d4d647a0f6910d80b94c55ad5a

Observation 8ef1edc8-d825-473e-9a00-c70a715886c3 · outbound

This paper cites Fine-tuning can cripple your foundation model; preserving features may be the solution.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Fine-tuning can cripple your foundation model; preserving features may be the solution

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.539321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.539321Z digest=sha256:656617686587e1f51a378fe3539fc5ecc2eaeebc163ed046bcd5dd14da2b6244

Observation fe1cb680-efc2-46f1-8f95-92bdad269973 · outbound

This paper cites GPT-4 Technical Report.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.661089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.661089Z digest=sha256:a82715e5f049026cdc8fb2d4ce77571f073175039c5d51105c0724f2089d85f2

Observation 2ea3413d-f36c-41d0-865e-7460a6dde878 · outbound

This paper cites Red Teaming Language Models with Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Red Teaming Language Models with Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.806561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.806561Z digest=sha256:f826fb46e70d2de2288e27e7e01280adae36ad465d023b35c6349a5c78a28b66

Observation 9220a8f9-c9e5-46fc-8905-9519f83138e7 · outbound

This paper cites Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:56.867587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:56.867587Z digest=sha256:b02215f468c6731e5ff9476a8539224bdb71f64b2f16549a5270e29969e9cca0

Observation cdf18a32-93eb-43b0-8018-ec1683e8b1d8 · outbound

This paper cites Model Extrapolation Expedites Alignment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Model Extrapolation Expedites Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.259827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.259827Z digest=sha256:c8219eda0d5a7821adca3a4fa1cbb328ccd662878b49ff467c29db3865cb8224

Observation e8bbe007-6169-48fa-8d86-030ac4ed2cb4 · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.357277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.357277Z digest=sha256:33eaf08e348188e6cad3b2a045d806f170c34cb674a8dcd9da3fa42e4ff1a73b

Observation d7e542b3-4915-4b0f-96d5-d0b1bd9ca74b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.417182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.417182Z digest=sha256:3950126c11787d0b0b92bb2e80e5619082f46ac6320f191a645ec277993af724

Observation 0661f49f-101a-41e5-9eca-3773a6fe15ef · outbound

This paper cites Similar to RIFT, RoAST can be categorized as a fine-tuning stage defense, and therefore is also not suitable for the scenario described.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Similar to RIFT, RoAST can be categorized as a fine-tuning stage defense, and therefore is also not suitable for the scenario described

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:02.053821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:57.532604Z digest=sha256:04c2ff55f985bb8d60c7019f97a191eee9a8472434f7a603a57dbaa8d4342f37

Observation 1d0977e8-b7cd-4f02-a3ef-b856ebc2bbdf · outbound

This paper cites Absolutely Obedient.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Absolutely Obedient

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:01.511263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:57.721181Z digest=sha256:a8483433a147080ba3db4f54a3dc797eee27bece762ba9617cab69034ae46938

Observation 5044b574-3f90-44e4-94a8-fc3683c08e2e · outbound

This paper cites Pure Bad.The Pure Bad dataset consists of 100 harmful examples, extracted from the Anthropic Red Teaming Dataset (Ganguli et al., 2022).

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Pure Bad.The Pure Bad dataset consists of 100 harmful examples, extracted from the Anthropic Red Teaming Dataset (Ganguli et al., 2022)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:01.311559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:57.801619Z digest=sha256:a13ac249ee211f9f2b694c80e660288be9881e2fa60f21263439f8420859316b

Observation 57d95072-760a-472a-9eb4-a7975e5adee5 · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:00:01.078057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:58.009728Z digest=sha256:e5ea45321195f4dfc1cdb8eb93bd2a14ba57235114bb811cf468c4aeac04df44

Observation d30f529e-ab95-46dc-a827-5232a6e9eb52 · outbound

This paper cites 5, and extrapolating further (which we considered broken), following (Lin et al., 2023).

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning 5, and extrapolating further (which we considered broken), following (Lin et al., 2023)

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:00:00.840892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:58.129917Z digest=sha256:dd2f009e4598a9b4c95359cc99f33f9445e00ed8e014a9fe972a73cca4f64403

Observation 74c87f94-dbf0-43da-9f53-33c61064ac15 · outbound

This paper cites We include Meta’s usage guidelines1 in our prompt, following the evaluation protocol of Qi et al.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning We include Meta’s usage guidelines1 in our prompt, following the evaluation protocol of Qi et al

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.388349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:58.303211Z digest=sha256:db976e00dc4cb577da81e1d2dc2a9c7ef3436b81e3f4cbb37e7f0a591c820a1f

Observation c9692a5e-b101-4f08-a28c-0d776e7391c4 · outbound

This paper cites No" Response(k=6,α=1.5): “Your task is to complete tasks for people is not recommended.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning No" Response(k=6,α=1.5): “Your task is to complete tasks for people is not recommended

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:00.202159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:58.413814Z digest=sha256:3f41d1e30e76fe44f1be030eeeaf20c61731d20e7a3b659caaf1d65a35eb7abe

Observation a00007b1-ff82-47a3-a550-603b29d4d067 · outbound

This paper cites an unresolved cited work.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Unresolved cited work

Reference 1024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:00:01.806107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:57.615080Z digest=sha256:f8220c41d3e735757ae886c706ad98c423ed7c3ce886f828a533288c4b7b11f4

Observation b71990e6-68f0-4e20-9c57-5b69d7972036 · outbound

This paper cites Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.134832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.134832Z digest=sha256:91a0ba63e1991c52c3c61923ce0bd3a8125ce4f18a162ce3966ceaf467208e66

Observation c3dd5955-0634-471c-906a-271f06b94965 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.560201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.560201Z digest=sha256:77e0f93fde03ddcbb77a9e6a15f85001c28993bf15e948bc0bc38f74df7437b4

Observation ed87a7e4-93cb-482a-976a-57b3b42fc5a7 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.503290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.503290Z digest=sha256:15a72c3b7b0a84634fcd40496387770a1bab8aae8fc9a1952f712184ca631030

Observation 96de15c3-91af-4fc2-8b70-29d0d129cb93 · outbound

This paper cites Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:00:02.261337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:59:55.714750Z digest=sha256:d3edfe9f4702f3eca16e7df47176f7bc9e3cc3fd8b670bebe9625aa9b4225092

Observation f948c392-9160-4e76-bb23-1f01999edd5b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:55.454095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:55.454095Z digest=sha256:d577388db8f30bc7390606815d15be8d43aed4c59929e4d52d6b7985fe83ad35

Observation ce426289-9b95-4344-a6ca-35b38b0d5789 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning Crowdsourcing Multiple Choice Science Questions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:57.027070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:57.027070Z digest=sha256:4b7346e7c0bd28dbcf3021d90161333a47b7ce4c3fa6d8214e43f76ae4ceb399

Pith citing papers

Observation 3b7f2f16-529f-46d7-b07b-d0fddd69fd39 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.060044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:1cfa30ac624208e1b50cab19351e7502559074f28694dbb66743b2ab76867e9b

Observation 12261f88-9c3a-4ae5-ac59-179bf9f182b1 · inbound

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints cites this paper.

Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:01.852959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:04:25.851592Z digest=sha256:0c06ea79c8f5c348532c4ba303175f97a18fe2c957232352f9cb0186a121bd73

Observation 27fd0429-61e2-4e22-b079-1ae9298dbdf0 · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:05:04.062857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:8485d1bfbfa0afdfe297a61ebf4e5b505823798cdbd265733f1a01e5321fb259