Pith. sign in

Paper Citation Record · LEDGER

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2605.10777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10777 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:26:13.390232Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact10
  • verified fuzzy37
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1146a942-e9e3-45b8-ac95-2ba6c4707be4 · outbound

This paper cites PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transforma- tion and Graph Compilation.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transforma- tion and Graph Compilation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:26:19.750035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:64c20f6ac98e67c0bdb124b46956949013dc0750056800bcb9d5eca5c6791f76

Observation 4a5cb2d8-04d3-441a-8c79-049b349aec06 · outbound

This paper cites Rezero is all you need: Fast convergence at large depth.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Rezero is all you need: Fast convergence at large depth

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.111699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:43f95bd3eaaef17236330a28a165e894a1fd384d660bb3bd07f993364ec1f56c

Observation 74e24c33-00d4-4b14-9b9a-8aff0953f4ea · outbound

This paper cites Considerations for governing open foundation models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Considerations for governing open foundation models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.205357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:46029774bfcdd7ffecbe776c8ad6fab9906cf85e45e4e51d7f97ba1c87ce61f4

Observation edafb6cf-e459-4b8c-be2d-4bdf07ced4c7 · outbound

This paper cites Convex optimization.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Convex optimization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.108414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:0fdfebfce97a193db818714b3357b53a035ff79444989ad868251f2344c52c07

Observation 7abbae8a-218a-41bf-bba6-22ca9e906757 · outbound

This paper cites Distillation Scaling Laws.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Distillation Scaling Laws

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:16:28.199224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:f889ae01e47ee68bf59a184f8dc55dc1ae9c4950e6c021563555bcecb68cc2f9

Observation 08bc4166-72ed-4fe7-93dd-ad84d01dcdd5 · outbound

This paper cites Open technical problems in open-weight AI model risk management.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Open technical problems in open-weight AI model risk management

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.115184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:e0edd5b430e189d019584ba6504d7c49b693ee2b9462d44191c12fdd7de0f5d7

Observation 1c8bfe88-8faa-44d2-bfb3-851ae2253786 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Training Deep Nets with Sublinear Memory Cost

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:28.211982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:e89aa94c781060e4e0734600e76b1eb2a54eaa1fb99d7d5434bc1fa4c1a5bab2

Observation bcb8e05c-6c3a-42a8-ad10-b21a967a7e5f · outbound

This paper cites Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:28.159412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:5ed872583b4a68cfffd89906f2bf5e3fb220a20f7a7c14a73e361a2014bf9961

Observation 1662ecce-cdd5-485a-a644-ef0746cef39b · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.177388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:646c107dae3e148605ba53e3cf5f3151ff0d0cb86062b8b7be80ed24dc171837

Observation 17d561dc-8334-4753-a7d1-c0c5aac3ebac · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Sigmoid-weighted linear units for neural network function approximation in reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.208526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:b52e89f06bc1f23993f9a74f2ade5b4bc38d879ea2809a6c7c00d520a6718958

Observation c907fbc3-a09d-4d07-936a-f324b5030233 · outbound

This paper cites Towards LLM unlearning resilient to relearning attacks: A sharpness-aware minimization perspective and beyond.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Towards LLM unlearning resilient to relearning attacks: A sharpness-aware minimization perspective and beyond

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.098223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:6b7bb0676088288959106580a01677206ae780c3057005c87eab74f43acd3831

Observation 3558afd1-82ac-4e6f-b93e-568da0125eae · outbound

This paper cites Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, and Asaf Yehudai.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, and Asaf Yehudai

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:16:28.190212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:cb710f20a80c66b353e2583e74ca01d9628369427ddd2ac5bf477342468532fb

Observation e3c25075-141d-4d56-9823-40fce105a36d · outbound

This paper cites On the symmetries of deep learning models and their internal representations.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation On the symmetries of deep learning models and their internal representations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.090827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:1707bf89ecbddadf58c2d5788014461422a0febc05f4a807c4aa6cddcdac13a0

Observation 32b4c2a5-2955-488b-8e91-4ccdf4033bd0 · outbound

This paper cites The Llama 3 Herd of Models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation The Llama 3 Herd of Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:28.204900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:001287fcbcd9b80c65dde3472f46539b479692feee69baa04c38d6f0df61dccd

Observation 036b333f-8477-44d5-893a-35ca6dd68bde · outbound

This paper cites Evaluating derivatives: principles and techniques of algorithmic differentiation.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Evaluating derivatives: principles and techniques of algorithmic differentiation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.187572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:eadfb8c97a747a4dd97b76c63d44995eba7f9465534dcbd43bbb8d24ed5f268f

Observation 10df6a98-4eee-44ad-9c44-c5ea99d7bd46 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Distilling the Knowledge in a Neural Network

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:28.182058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:ed700d246f7f88e080d525f06bcd8fd073decf68973b1cd6117e0624b7357184

Observation ab32ee82-f0c0-4bf8-adda-080e6747aff8 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Lo RA : Low-rank adaptation of large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.063760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:a3e7b6255bd3d353adc7258a02885566dc4ffeaab2825b68e080b983fa5f0406

Observation 9d23481f-d0f2-4b60-85ee-cf6803769191 · outbound

This paper cites Unlearning or obfuscating? jogging the memory of unlearned LLM s via benign relearning.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Unlearning or obfuscating? jogging the memory of unlearned LLM s via benign relearning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.101413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:acd4b7d6f7880642299abd708ca90d3814efd58a7ab2a3b1b346d69bf2dc5e58

Observation 52e29308-b72e-4f3a-b007-bfea6bcd1d7d · outbound

This paper cites Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Vaccine: Perturbation-aware alignment for large language models against harmful fine-tuning attack

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.086812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:b8035b4dc31c7e3c4c48008ab513a689dc52b0c948f68f39cb85198ab6d83a2e

Observation 817a1468-5c41-4194-b5a6-a28c76e74035 · outbound

This paper cites Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Booster: Tackling harmful fine-tuning for large language models via attenuating harmful perturbation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.154570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:36b8134a926510446a08ae5d9f26f99b03bf8b83caa7b4b604b651cc4fb181b7

Observation 7b8976ec-21f3-4f33-a5e4-21bc2c052db8 · outbound

This paper cites Hutchinson.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Hutchinson

Reference 21

Resolution
verified exact
doi, observed 2026-05-12T04:26:19.761188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:32d8911a422c25841887f6fa642d628abb937e851c2f85348f17321b9d9e0aba

Observation 5baa2546-bafb-4ac4-99da-c92178d2684f · outbound

This paper cites Disrupting model merging: A parameter-level defense without sacrificing accuracy.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Disrupting model merging: A parameter-level defense without sacrificing accuracy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.119641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:23646b54524b1b9500f1eaaa39acf9ad0a538e10f1f5bc73e9f3403f7b9e5ba1

Observation 19fef098-1275-4281-9259-800b95c4c31e · outbound

This paper cites On the Societal Impact of Open Foundation Models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation On the Societal Impact of Open Foundation Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:16:28.174858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:736843c5b885a15d2ea1013316edd2351bc69662a40a5ce8caae5797d6ea574d

Observation 12794b98-64b9-4fba-8bcb-54cb807868a2 · outbound

This paper cites La cryptographie militaire.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation La cryptographie militaire

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.104885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:1e53d3bd1b08f2c1d27788c3e454883bb1060e248527be32115d677e48c00dd8

Observation d7a5f948-d4c9-49cf-a558-33d8ce3edbd6 · outbound

This paper cites Mnist handwritten digit database.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Mnist handwritten digit database

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.094475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:7eab6017cb9b5140c8829ebbfbb0f5f1d52179622a9ec88a17c820a3f5433b23

Observation fd5c7b35-e46c-4543-b0a4-76fc5f5528f7 · outbound

This paper cites Lee, Addie Foote, Alex Infanger, Leni Shor, Harish K Kamath, Jacob Goldman-Wetzler, Bryce Woodworth, Alex Cloud, and Alexander Matt Turner.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Lee, Addie Foote, Alex Infanger, Leni Shor, Harish K Kamath, Jacob Goldman-Wetzler, Bryce Woodworth, Alex Cloud, and Alexander Matt Turner

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.183929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:1b00e45d3fc83baac045da1816b5978b36d6ff56d26b4aef8af1d0db9a39961e

Observation f0f564eb-d8b4-4221-bc04-9f6c048fc86d · outbound

This paper cites Module-wise adaptive distillation for multimodality foundation models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Module-wise adaptive distillation for multimodality foundation models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.198609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:e92bf46c077ba7ef092dc6b2df4d649e584215050fa2777c20c370d700c292e9

Observation 64b30074-c4ba-4ec9-95ee-04043a710959 · outbound

This paper cites Less is more: Task-aware layer-wise distillation for language model compression.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Less is more: Task-aware layer-wise distillation for language model compression

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.201757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:2a76a441c5480379ade0c6074170fb5b2079616169cdf225f81aeac0caab33ed

Observation cf9d5e9d-c275-4e6e-9457-5d6f8d3296dd · outbound

This paper cites Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Targeted vaccine: Safety alignment for large language models against harmful fine-tuning via layer-wise perturbation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.191273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:3821cb72712aafcc1f298f4a4580080c687b6b1a9580eb94df24db06c977f231

Observation 6438e206-5615-4d87-8bf8-bdd7e8b4453d · outbound

This paper cites m2mKD: Module-to-Module Knowledge Distillation for Modular Transformers.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation m2mKD: Module-to-Module Knowledge Distillation for Modular Transformers

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:28.220035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:f3178dd24ec522f9a9c9a199d47ba3a8344af5f426e6c25b2897ae6de2cc38bb

Observation 8aca8dd0-e0a0-4d17-9116-5a9d54493c2c · outbound

This paper cites Decoupled weight decay regularization.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Decoupled weight decay regularization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.137535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:8f2c46c00a207f3c3eb192c199ae47c0b108c60aaeb02932305dbca538703016

Observation 224771bb-623b-4a19-b09b-7238f659527b · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:16:28.240944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:383bb6da069ec28a5309a264cd10f1b606bbb032992f21c40a76a7e4c51236e9

Observation 67b7680f-a3a2-4c65-a733-78a8c574347c · outbound

This paper cites Pointer sentinel mixture models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Pointer sentinel mixture models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.148395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:eb9b6979a7adedfe7045135828b79d1acd547872bdab87990ece2317994a756c

Observation ea66fccc-e7b5-48f4-9419-f1a9999126b5 · outbound

This paper cites Antibody: Strengthening defense against harmful fine-tuning for large language models via attenuating harmful gradient influence.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Antibody: Strengthening defense against harmful fine-tuning for large language models via attenuating harmful gradient influence

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.151782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:59dbe626e90edd4ee5b1a3639d08758b1a8d975d9216d973f56ebcbe45302049

Observation 3e92ec92-0f11-46da-b1e7-295b1e156269 · outbound

This paper cites On evaluating the durability of safeguards for open-weight LLM s.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation On evaluating the durability of safeguards for open-weight LLM s

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.141460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:315c90658653905c04a22c773e799092a411a3fabc574882b80e2208d8acae7e

Observation 06ccddfc-3491-472a-9113-488e9886a119 · outbound

This paper cites On-the-fly adaptive distillation of transformer to dual-state linear attention for long-context LLM serving.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation On-the-fly adaptive distillation of transformer to dual-state linear attention for long-context LLM serving

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.170574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:0706b22fd63250a0ff7fb0d295165178892d571c8f4408416c43b9cf4930eed5

Observation b096501d-cbfd-442a-b274-3bd087e188d6 · outbound

This paper cites Representation noising: A defence mechanism against harmful finetuning.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Representation noising: A defence mechanism against harmful finetuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.057592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:c8c52d50f9187f1d59de0c82d66349e90692f81c9e2827a925df7118b161ed38

Observation 6fa48836-8e29-45ac-b07d-e5ea1fbe5267 · outbound

This paper cites Locking open weight models with spectral deformation.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Locking open weight models with spectral deformation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.048013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:b061ce8d03597d99332349a88714c7705f2bea79a89cd90e74311beb75a0b07a

Observation 6ec0bf43-9976-4023-971a-671e7b05e4bc · outbound

This paper cites Limits of convergence-rate control for open-weight safety.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Limits of convergence-rate control for open-weight safety

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:28.227937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:cda77ec855e9f17290ba583bdcab852382073883492891a094cf54d02f22bfcc

Observation 5b971fa3-3348-4d8e-9740-352b7a29a59a · outbound

This paper cites GLU Variants Improve Transformer.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation GLU Variants Improve Transformer

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:28.233948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:39a0221558728f9e34184be703397cbb87f936ef6820a0a8f7740c1d3c50ca66

Observation 6a1bf5bc-f230-46fb-a202-fef2a15c569d · outbound

This paper cites Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.194644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:9d2aeedfa2bf0e9d2484b1e6db3ac37efb2b615b2425c9998fe684dbdd813932

Observation e52b4358-e8af-4b05-be7b-c29abc9a7539 · outbound

This paper cites Tamper-resistant safeguards for open-weight LLM s.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Tamper-resistant safeguards for open-weight LLM s

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.166944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:d5ad997986cc76a902d19ab7daac38968b481a80e8c79297f572c3dfc88553b9

Observation eeaf311c-e4d2-4ed8-94c1-c1fb5f3b9144 · outbound

This paper cites Attention is all you need.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Attention is all you need

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.174080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:b28080890578beed98209dcfdf8e81cee3ba9ef0f0c71b60f24efa36a2fc2804

Observation 775eccb7-86f8-4244-a26e-a22d43c0aedb · outbound

This paper cites Self-destructive language models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Self-destructive language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.145128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:149bc647633e7c3f0048cae86e6f08adc38496dba87aaeef490871480b05e442

Observation 73fd0f3e-ed64-4a4b-af11-c857908ace5e · outbound

This paper cites Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:16:28.247059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:21d2b476ff31b80c9af8f467081ebb3002a6a6d978fa91ebbba456c9604c3b8a

Observation 53ccb8b3-3a47-4ed8-b451-7a967801d77c · outbound

This paper cites Towards building non-fine-tunable foundation models.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Towards building non-fine-tunable foundation models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:28.255212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:124afd924939f8f4d74a47caae9969d711067ad1716d6241ad7c179ff6209150

Observation 34d8b140-e1de-4225-92e7-a0a6ead138df · outbound

This paper cites Bert-of-theseus: Compressing bert by progressive module replacing.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Bert-of-theseus: Compressing bert by progressive module replacing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.158246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:6db67aba89a2f4feb578729880dd8270d5860185a26061e6dac9e973591d102d

Observation ae51a38c-172b-4949-ac0e-0577a09820e4 · outbound

This paper cites Qwen3 Technical Report.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Qwen3 Technical Report

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:28.165173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:7405d82672ce6987e78c0575c9bf991f7c4f19e29393f45d52621e608f52c794

Observation a11eab7c-583f-4846-b6d5-67e096e72acd · outbound

This paper cites Asft: Anchoring safety during llm fine-tuning within narrow safety basin.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Asft: Anchoring safety during llm fine-tuning within narrow safety basin

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.163682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:cf8f9a6f6ce8da6a98a81362c9cdf563587ca7434a31b078231597dfef082814

Observation c6909493-e337-4cd3-aab5-2a57efc76515 · outbound

This paper cites Root mean square layer normalization.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Root mean square layer normalization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.180418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:6b0592da7ff87b7bfcca4e78470e42395e21f7b8d60f482a002dcd0c7e36b647

Observation 36e56fea-88ca-483d-b611-75c1acb17003 · outbound

This paper cites Symmetry in neural network parameter spaces.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Symmetry in neural network parameter spaces

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.053925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:b4fbc3deda107e0c2ef429ed761da1856e463aab26ec2bb0efc40d6bef011341

Observation 9904225a-f0bc-410a-9a31-3da0593dcb55 · outbound

This paper cites Understanding and enhancing safety mechanisms of LLM s via safety-specific neuron.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Understanding and enhancing safety mechanisms of LLM s via safety-specific neuron

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.060606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:802d0b5f71eced49a3d504efcd0012d84a35e9528a804762fe5b635f24e985dd

Observation 62e517c7-d354-4fdf-ab4c-a9d0daac584b · outbound

This paper cites an unresolved cited work.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-12T15:06:39.160723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:342bf2315030d5ae506528895fe39c17c304559ccf58c008cef59bfe022e5239

Observation ff30451e-5da0-476b-b886-66b1eb42b9a7 · outbound

This paper cites Modular transformers: Compressing transformers into modularized layers for flexible efficient inference.

Locking Pretrained Weights via Deep Low-Rank Residual Distillation Modular transformers: Compressing transformers into modularized layers for flexible efficient inference

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T15:06:39.123334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T04:26:13.390232Z digest=sha256:8c33e11f84116b733dda431b00c4254d689d725b0aa379d63945dc2e7c9f9d80

Pith citing papers

No inbound Pith citation observations are available.