Pith. sign in

Paper Citation Record · LEDGER

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

As of 23 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2502.00840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00840 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:39.675692Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:00:19.630547Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T23:10:21.867778Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2140333-e41d-4262-babf-344189700f7e · outbound

This paper cites Chat generative pre-trained transformer (chat- gpt).

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Chat generative pre-trained transformer (chat- gpt)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.445405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.445405Z digest=sha256:ef2139a4161f971a478817e6131299a5a65b724b7588dff0347bb2669f007fc1

Observation 0cea145c-5b3c-45de-a31e-bd36475a504c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.449262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.449262Z digest=sha256:0e44cbd6ade0e2fa8e9ab9aaa7286b91fc706eb670c95f203cd20a5ce11bc453

Observation 5ba827a8-8fac-4239-962e-599a67040617 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gemma: Open Models Based on Gemini Research and Technology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.452541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.452541Z digest=sha256:18e5165dc7d024629e275476497a69f3979540f9b462943c05ed15a535134325

Observation 74610e66-12fc-4343-8437-0fcfe916ef52 · outbound

This paper cites Mistral 7B.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mistral 7B

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.455573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.455573Z digest=sha256:eca487eecd8bfdfceeb15a942c3181cfb562b3130798659005b45d41eab62ea7

Observation 01e4ab5a-08e7-4396-9012-3dfd8245a74f · outbound

This paper cites The Falcon Series of Open Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense The Falcon Series of Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.458365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.458365Z digest=sha256:bf5a7dd09c47229fa48977fe4f85ff0bb2630ef99219a7f5c64d35600731d817

Observation 64401fd9-c23c-4d59-aa32-5ca14ea0d8f3 · outbound

This paper cites Qwen Technical Report.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.461305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.461305Z digest=sha256:441033845c62a076152de96840c3e5fa47cbcfc7ae558db839cfb90291b4c5f9

Observation d0ec0fee-50b6-4ebe-92af-222ca99d8f52 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Gaussian Error Linear Units (GELUs)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.464411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.464411Z digest=sha256:b7514861c23f45326100f50dfac0d64c98c6ab4db5e93be4d48d2aa0b928955d

Observation c704735d-6f10-4dcb-b2a8-67c736845040 · outbound

This paper cites Sigmoid- weighted linear units for neural network function ap- proximation in reinforcement learning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sigmoid- weighted linear units for neural network function ap- proximation in reinforcement learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.467467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.467467Z digest=sha256:53176277cedd79ff6f9f0973055e723adb65a609dc71ead843cbe05d7ca7126b

Observation a6aaae0c-d6cd-4a15-bb84-41a7bbc6a0cb · outbound

This paper cites Awq: Activation- aware weight quantization for on-device llm compres- sion and acceleration.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Awq: Activation- aware weight quantization for on-device llm compres- sion and acceleration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.470224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.470224Z digest=sha256:01d74c6adadfc2b19da5e97694adba25c2d04d0adac42d314776b37a603eaee4

Observation 35836eb3-c6d7-437f-8d81-84cfea072082 · outbound

This paper cites GPTQ: Accurate post-training compression for generative pretrained transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate post-training compression for generative pretrained transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.472818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.472818Z digest=sha256:4347cab6ea6cdbbf6f285cf2014d6dcd068432277fb4e20995c4138f5336b7c1

Observation 479f9a15-8326-4792-8bfd-ec4a0f8371c4 · outbound

This paper cites Billm: Pushing the limit of post-training quan- tization for llms.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Billm: Pushing the limit of post-training quan- tization for llms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.475289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.475289Z digest=sha256:18cd6d869a0f90b86b430132716ce5a293e8365adcd705a20a77b7c7ea84f3fa

Observation 651d0dfb-4ab2-45ee-9c83-d4503e1b4c98 · outbound

This paper cites Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.477774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.477774Z digest=sha256:8a77fcc037dcd544245f97713d62f3440efc4e2b884570b48f7c7e1ae3a47706

Observation 2bbb1169-9b05-45db-bc95-7c1ca97e02af · outbound

This paper cites Llm- pruner: On the structural pruning of large language mod- els.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Llm- pruner: On the structural pruning of large language mod- els

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.480322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.480322Z digest=sha256:eaa0127b699afd06db44299ed0001b26075df3855e2f7e3d587af4a82a9036f5

Observation d7a430ef-fed4-41b6-b1ba-6de630ba7c3e · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.482768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.482768Z digest=sha256:35e9c4940497501cce66e5318b69918ece74d981e750023380f1a80189fa0047

Observation faed12e6-880d-4a75-9f5a-8bbbb02903de · outbound

This paper cites Structured Pruning of Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Structured Pruning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.485481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.485481Z digest=sha256:93a54071d4293b368285f31ae68ee6cb1f55424de9f82aec3958302dbfc14b96

Observation 058c99be-1b39-4bf5-9a2b-a218ac461f99 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fast inference from transformers via speculative decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.488230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.488230Z digest=sha256:01920ac3d48e6c76a73d95b7ff1447e14857025d3d60b940f7bc58e49a3cb2b5

Observation 68d3aff6-70f9-42cc-b2b5-c48b44aca84b · outbound

This paper cites Speculative decoding with big little de- coder.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Speculative decoding with big little de- coder

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.490567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.490567Z digest=sha256:8fbc069386a142d3f0857fe31358fea876dc3b6bfa206ff47e3e1ad6a71ba9f2

Observation bdda5e25-d4fe-4fb3-92b7-8fe140a74b12 · outbound

This paper cites Iron: Private infer- ence on transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Iron: Private infer- ence on transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.492867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.492867Z digest=sha256:554d9d773e134e5a362c316ccbbc97e1484f29a8e8fbcebf96c97e2c0fc73777

Observation a5dafbbe-a6ca-40c3-a41c-8fb73b586f97 · outbound

This paper cites Ciphergpt: Secure two- party gpt inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Ciphergpt: Secure two- party gpt inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.494884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.494884Z digest=sha256:f4752055fedefb5a66f0e1fedbb2a24693c429eb1c3ebbdbc3f8ed5ff31aa8ad

Observation be552ba8-a521-4ba9-9f1b-57b70edcef59 · outbound

This paper cites Bolt: Privacy-preserving, accu- rate and efficient inference for transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Bolt: Privacy-preserving, accu- rate and efficient inference for transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.496815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.496815Z digest=sha256:3c26e288799fd5a9c9f0812cc1c21d8f6165f12ad9bd4962a143eb10cc98079c

Observation 19367cde-fe3b-4f0e-b017-bbd0653347c2 · outbound

This paper cites BumbleBee: Secure Two-party Inference Frame- work for Large Transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense BumbleBee: Secure Two-party Inference Frame- work for Large Transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.208924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.498787Z digest=sha256:8fe1db1552dd04efc65c5a111ae99f7294f1770becd01d8478cdbce5f1550415

Observation 1b0a96f9-7cf6-43b7-8d4a-bccdc494eec1 · outbound

This paper cites Secure transformer infer- ence made non-interactive.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Secure transformer infer- ence made non-interactive

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.194371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.502964Z digest=sha256:9d0c7b4be3e2664f7c77a7f37e7205fb8a16bb9402db8b8b1ae6e97a0aea854c

Observation 4e6321c5-dda7-4236-96aa-7c40970326c2 · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training-Free Activation Sparsity in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.505073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.505073Z digest=sha256:a2995815eb13657ddd9844fcc9661d1d375eff8af40068be744b6db114ef41a0

Observation 287bf1e2-4352-430b-ac76-5ab02d64ee48 · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.507833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.507833Z digest=sha256:c5f8384c2d7b8ae3b19beafa20f664515bd47fa212f317c8355a01e3cbaf640d

Observation dcc82b43-ba2b-4f9a-b7d9-6cbb855b197e · outbound

This paper cites Relu strikes back: Exploiting activation sparsity in large lan- guage models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Relu strikes back: Exploiting activation sparsity in large lan- guage models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.187461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.510440Z digest=sha256:6e6e789d4d03346139657cf1a14c66fd0a00d57786a407ee97f224f901b2c40c

Observation 288148a7-7b86-42e1-af4f-39648d09b0d8 · outbound

This paper cites Smoothquant: Accu- rate and efficient post-training quantization for large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Smoothquant: Accu- rate and efficient post-training quantization for large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.180020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.512941Z digest=sha256:419f77e9ded7ad4e9d79835070749fdbe4bec13abeb952ab218a7b4e373cdc25

Observation a1d2126a-2aa5-4d7a-869f-5e4a87678d1a · outbound

This paper cites Omniquant: Omnidirection- ally calibrated quantization for large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Omniquant: Omnidirection- ally calibrated quantization for large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.172593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.515463Z digest=sha256:b1bd64dbe6fc713bff6d9689bb92ba2b5bdd096412f31217f49c90676c1f7b63

Observation 8c840334-b818-4007-91e1-3b4854e5963b · outbound

This paper cites Sirnn: A math library for secure rnn inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sirnn: A math library for secure rnn inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.165223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.517951Z digest=sha256:56efb39ddf0e10d834e58dcb2bd7bcac7eaa004934e1ae7eb4a26b0134d83d6e

Observation 4f1e32b9-8c20-42b2-84ef-e4f526db4df2 · outbound

This paper cites From chatgpt to threat- gpt: Impact of generative ai in cybersecurity and privacy.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense From chatgpt to threat- gpt: Impact of generative ai in cybersecurity and privacy

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.157872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.520323Z digest=sha256:3c41864ab160c382fa39ad4bd2304cac647860adbf8e397d72d7ae924703c33f

Observation 637b992f-b3f1-45cf-ba58-87c2879cdb14 · outbound

This paper cites Exploiting large language models (llms) through decep- tion techniques and persuasion principles.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting large language models (llms) through decep- tion techniques and persuasion principles

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.150589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.522839Z digest=sha256:ba899b968f4403f75ef82d58d001aa4ffc8c3c7ef61e2c539a62314a325fa1e1

Observation 4e4e6dbc-4e9c-4b45-afbe-ce3af4e2d85b · outbound

This paper cites PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.525244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.525244Z digest=sha256:aa05ade3a25eb3f2566923d894fea2da50be097ca12219396d14a76e58af469a

Observation 1ef5ecec-76c2-49c2-9151-df6e15127c28 · outbound

This paper cites Training language models to follow instructions with human feedback.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.527855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.527855Z digest=sha256:938150414172b6ca180eed2063afc4117249328d26d93ebda8d4d180a7eecd06

Observation c76fa3c1-23e2-4af4-95c2-8bb3a4510ec2 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.530426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.530426Z digest=sha256:96aaa1f445f8a39cbbcd5ad0b9249fa80a7acc2fd01f780cebee5f4b3c9c45b4

Observation 990fe8fb-ecd8-4a21-a561-8001f009b155 · outbound

This paper cites Di- rect preference optimization: Your language model is secretly a reward model.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Di- rect preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.139508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.533289Z digest=sha256:d0464bda44c11cee78574ba0462bbe77418e9e8c375345ca47ea021ae9fb8808

Observation 28292dea-c199-4d8d-b0c6-068dbfc45920 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.535809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.535809Z digest=sha256:f82304448d585e3ef4f27739cb55db84a4d7fa9689e3a536c905689f571204dc

Observation 21dcb9ae-d6d4-444f-96b0-4f402b7abfbe · outbound

This paper cites Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.538531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.538531Z digest=sha256:4d7ac3664e9ef61969d8ddf9c70a17b881c5e4406c9a9e920c878ca326494e9f

Observation 648e5c71-b1cd-46a7-b4ce-4f903832b0df · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.541244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.541244Z digest=sha256:693fb2171e7f1a3af2d8f582703d00aac1c54ead25c27d61d04da8301b01d060

Observation 554ca386-87ae-4910-a4bb-96a9a97df6b0 · outbound

This paper cites A simple and effective pruning approach for large lan- guage models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense A simple and effective pruning approach for large lan- guage models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.132421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.544127Z digest=sha256:e125ccb454ee8c69cb62321ee03270435775eff7aeaf75f4bd97d41a4db81d75

Observation eb45a763-0ab2-44c5-be77-10d227b55ad4 · outbound

This paper cites Dynamic context pruning for efficient and interpretable autoregressive transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Dynamic context pruning for efficient and interpretable autoregressive transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.125481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.546825Z digest=sha256:20ae080acb8d580b946ddca8e674f96fff7e307ba30538e17998f1b8c9e60641

Observation 7ce3b523-7318-4acf-8037-686931756f86 · outbound

This paper cites Adapting language models to compress contexts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Adapting language models to compress contexts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.118461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.549206Z digest=sha256:4495c7be6504241389d9bab506a7f337e32a04b425149fe0e321612bc66e4c07

Observation c805132b-81c1-4aa9-86dc-05ae00074791 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {InfiniGen}: Efficient generative inference of large language models with dynamic {KV} cache manage- ment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.111328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.551730Z digest=sha256:8a2d5a816ff79840ee25adad9e9242f4d5f9da4cea0e68d79414a420e1f161a9

Observation 71451ffd-5c83-4a1e-a185-f5420ad3677c · outbound

This paper cites {IMPRESS}: An {Importance- Informed}{Multi-Tier} prefix {KV} storage system for large language model inference.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense {IMPRESS}: An {Importance- Informed}{Multi-Tier} prefix {KV} storage system for large language model inference

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.104171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.554120Z digest=sha256:5cd628125185500056a63e8646c66a58a9af222768cbb8127b7930fa260c6f83

Observation 3a1c35c4-aa16-4794-9217-f105d484939a · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.098118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.556605Z digest=sha256:31fff95a167190759fd4c0e58d3821bcab6213b472bd559c28d736a87771e178

Observation e7788879-bce4-4f23-b5e5-cc32f49dd2cb · outbound

This paper cites Kvquant: Towards 10 million con- text length llm inference with kv cache quantization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Kvquant: Towards 10 million con- text length llm inference with kv cache quantization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.091674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.559046Z digest=sha256:20efecdcf0ebeee34e6554d0e68ac30e11ed75b92e0e8fddf7a01ad515d96d56

Observation 10aad630-7787-44c9-855c-ac005805c07b · outbound

This paper cites Sequoia: Scalable and robust speculative de- coding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Sequoia: Scalable and robust speculative de- coding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.085231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.561483Z digest=sha256:34a3c8ba8fa52d7b0fc1553149a7ea54ffce724f2d50d3166109da6340c4bf4e

Observation 48c61911-7bba-476a-bea3-44a973b2df16 · outbound

This paper cites Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Decoding compressed trust: Scrutinizing the trustworthiness of efficient llms under compression

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.078252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.563904Z digest=sha256:ff167d3729ddbee284b6c12a95e9910b966fd23fe4eb18a4ab30926104f7db79

Observation 45a28834-227a-4d5b-8201-4e64193306de · outbound

This paper cites Harmlevelbench: Evaluating harm-level compliance and the impact of quantization on model alignment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmlevelbench: Evaluating harm-level compliance and the impact of quantization on model alignment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.070834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.566243Z digest=sha256:e6f698cda4eb374ffa43eb918be6307fb70af8d1432a1bd6d6fca254df58eae6

Observation 5b8a9931-e89a-415e-892e-742fb4724cb2 · outbound

This paper cites Exploiting llm quantization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploiting llm quantization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.063577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.568667Z digest=sha256:01b70e567e56dce67db8d0b98f0213e55d17d5f8f2f8c78fcd3aa008dbde7bfd

Observation ceffec12-a271-4d19-b0cb-f420f35f503f · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.056232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.571296Z digest=sha256:96dc2984c6c60f5fe282cfd5e61091147a4452ba01a14dc5f985d3ff89181382

Observation 58b43f76-c1a2-4ce1-b0e4-34a321323825 · outbound

This paper cites Qllm: Accurate and efficient low-bitwidth quantization for large language models.In- ternational Conference on Machine Learning, 2024.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qllm: Accurate and efficient low-bitwidth quantization for large language models.In- ternational Conference on Machine Learning, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.048307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.573726Z digest=sha256:1f85dba0f7f58ae2c8835e10aa24a76ee26ddfc4503ee7ac7f91bcb47b1ec9dd

Observation 4b9e6550-0094-4b81-ad21-6164bfb5beee · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Introducing meta llama 3: The most capable openly available llm to date

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.575830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.575830Z digest=sha256:48c022946b889011bb3ed30f7e87483aba4262f4788c0a1a7d46a283be5b48b5

Observation 036314c3-df4a-428e-85ef-917363701540 · outbound

This paper cites Mpcformer: fast, performant and private transformer inference with mpc.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mpcformer: fast, performant and private transformer inference with mpc

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.037095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.577796Z digest=sha256:43d4089ba4e6a0ac8961f33ed4b7f004e72983924a95a1c0ba4e9b78fd9afeec

Observation bdb1cc3d-e089-4a1e-9e20-bf96c91f299b · outbound

This paper cites Encryption-Friendly LLM Architecture.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Encryption-Friendly LLM Architecture

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.579915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.579915Z digest=sha256:1b49b1d78a85d0e97aea85d0f5afd119121db0d20206b69bcc94601c08bed665

Observation 9f2ac483-69c7-42ab-bbfb-31c9b162cb94 · outbound

This paper cites Deja vu: Con- textual sparsity for efficient llms at inference time.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Deja vu: Con- textual sparsity for efficient llms at inference time

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.029642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.582055Z digest=sha256:32dd057b14fb778c11098f26d8f51eb8a3862f95f2954082e22772bcdea490a3

Observation 6da7ef52-b474-4671-86df-2cb90b6c996c · outbound

This paper cites Cats: Context-aware thresh- olding for sparsity in large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Cats: Context-aware thresh- olding for sparsity in large language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.022246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.584000Z digest=sha256:5099a276c92cb9358f92d2a721602f98d13815c027ff23b8da04accc52680d8d

Observation d9906681-c789-4184-9a25-d9900b8443bc · outbound

This paper cites Judging llm-as- a-judge with mt-bench and chatbot arena.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Judging llm-as- a-judge with mt-bench and chatbot arena

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.586311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.586311Z digest=sha256:60aa51ef26a65a196e8b0bcf606195c835a3b00f9756af5e62fb921049e031c8

Observation 04a0dd48-3eb6-413b-94ef-63469149a2a4 · outbound

This paper cites Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Assessing the brittleness of safety alignment via pruning and low-rank modifica- tions

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.010797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.588409Z digest=sha256:e238d2efa229c67f10b2bab93a7cf0759fd5807f0a2cadfe0e1b76124369a0c3

Observation f9eaa58c-f081-48d1-9de3-5221b7b1162b · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.590442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.590442Z digest=sha256:2d9623c0522c5f9f06a29a59493dfabb07b2e6cc4f49eaf9c306c014719ca901

Observation b32f45a2-16e6-49de-8886-851d2c32f982 · outbound

This paper cites Lisa: Lazy safety alignment for large language models against harmful fine-tuning at- tack.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Lisa: Lazy safety alignment for large language models against harmful fine-tuning at- tack

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:40.003409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.593006Z digest=sha256:193483ea82d07a3d83195e703fb6a936f2177538bd275ba0c6940e5295d0dadc

Observation a5fb86ba-5cd5-4115-adb7-184980feab14 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.595425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.595425Z digest=sha256:4c4b18b052fecddb7a41a7239ec229acbcbc22a749521943c7e61420fdb0865b

Observation a2012904-6ca2-4b52-8886-23136f88c45a · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.598235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.598235Z digest=sha256:8a28ee621c784fcece0f5dbe26d08edeb2f8aa9f47e3e0ef1160c872e1b54aa8

Observation e98b3c49-8494-4c98-9d5a-0b2b8460beac · outbound

This paper cites Detoxifying Large Language Models via Knowledge Editing.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Detoxifying Large Language Models via Knowledge Editing

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.600920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.600920Z digest=sha256:92134f51a933326a279baa1497ea42c05512f7e55b5ba20a789fcc087025cae9

Observation 9ed5edc6-4186-4910-8e00-71df8537532c · outbound

This paper cites Recovering the Pre-Fine-Tuning Weights of Generative Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Recovering the Pre-Fine-Tuning Weights of Generative Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:39.818906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.603681Z digest=sha256:845e167cc9534afe785f4e85d7d6de9fea27d8f0d4db7e42652599b83fdc12a6

Observation dcecaae4-378e-4c13-a3a7-3cc5ef393661 · outbound

This paper cites Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.606386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.606386Z digest=sha256:3534a327f33559c7703bbe08b7b7819483b9b56946c7460eb97f1a431f5d2cb5

Observation 0d6d80ec-20af-4486-bbc5-73332f878ca0 · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.609496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.609496Z digest=sha256:9e72617a6d09df3f85cc32c2819ea85e66e7ab8526e49c7bc6538628f14a193d

Observation c1a338de-b2c4-4f79-8bcc-dacded386d85 · outbound

This paper cites SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.612801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.612801Z digest=sha256:05ebb1f9261ed4453ed074d56098afb522ed08511a7d52f8a357bc4323e2f8d3

Observation 6a15da23-7dd0-4138-8b6c-b2475b3ffc78 · outbound

This paper cites Compressing LLMs: The Truth is Rarely Pure and Never Simple.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Compressing LLMs: The Truth is Rarely Pure and Never Simple

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.615850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.615850Z digest=sha256:46c2532d2f0433d4303966bfda66432b51a57b978ce8ee4793ad9ef1d4933a8b

Observation 960c1af9-ac41-4784-877c-0b3489ae88f3 · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.618406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.618406Z digest=sha256:92d9fa4b298e8fbadb0bf2d57daf2eb66da3e787f6dd0ebad1d2ec2ebe3da07e

Observation d45dd858-0691-4600-b9dd-4eca0d8c4165 · outbound

This paper cites Discovering sparsity allo- cation for layer-wise pruning of large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Discovering sparsity allo- cation for layer-wise pruning of large language models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.996108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.620977Z digest=sha256:9c416459c40652b830fd5bd19460c044465a0e9a77381beb5aea9b529fde1c11

Observation 63d92d98-cf18-45a9-a45d-6a71f0bcdb09 · outbound

This paper cites Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.623449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.623449Z digest=sha256:4a1fe53064f03a5c571772448e68f691cc1c09e19b44f0575db7c4c662e282f5

Observation a13e492f-863c-42fb-816e-a0528327c55a · outbound

This paper cites Navigating Extremes: Dynamic Sparsity in Large Output Spaces.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Navigating Extremes: Dynamic Sparsity in Large Output Spaces

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-09T17:37:39.772417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.626079Z digest=sha256:c00738afe95323dcdac702ef08d47ced85562bd22d2d2f67dcbaf39e4f51d955

Observation 840f2a7a-b166-4dbd-8102-4b404d28ec5f · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.628959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.628959Z digest=sha256:67c3325b33e378e66ccac8084f0157c8245c8656a2bc780a5e2d29a9966d242a

Observation 67235692-f973-4818-a960-1f910571583d · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense TrustLLM: Trustworthiness in Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.631778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.631778Z digest=sha256:11981360f60943854405aab67bbca1f89d78b842890fa3e4229d5646d59e359d

Observation aa89ebd6-d450-44d1-9013-7aa020b7f05f · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.634473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.634473Z digest=sha256:aa6c63ab3f9fd8a0ab99eb5690b5f8f6b8f5752d5ea096ad68faff79d41eb10a

Observation 778cdd0d-f1e7-4867-b52e-1cc7557956df · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.988585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.637030Z digest=sha256:7ef4cac3a8e7970050b3b8c16cca61451f6f7d9d469b52566a31f47807a00982

Observation 7cdadd0f-7f26-40e0-83d7-617cc7654d8e · outbound

This paper cites Automatically auditing large language mod- els via discrete optimization.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Automatically auditing large language mod- els via discrete optimization

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.981980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.639157Z digest=sha256:28e5eb99a0d3536c542cc56acbe819627d25bf5bb21742ed6e350a8504eccd9a

Observation 894c6b5b-92f3-4d3d-9ee8-b8e810b3b87f · outbound

This paper cites Making them ask and an- swer: Jailbreaking large language models in few queries via disguise and reconstruction.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Making them ask and an- swer: Jailbreaking large language models in few queries via disguise and reconstruction

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.975523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.641111Z digest=sha256:6dc1fbef349e1d7f6fec5dae3f605b702a024bff5220f3341186fc4571b7554c

Observation 5df229dc-92c6-4790-b0ac-a7dcdadbf21b · outbound

This paper cites Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.968993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.643119Z digest=sha256:9f611322d323120a175b7b23eaaff5dfa42aeac235022d79a84083eca4830efb

Observation 334cc4f1-54fd-45de-949c-8bb81b5e6779 · outbound

This paper cites Pointer Sentinel Mixture Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Pointer Sentinel Mixture Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.645158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.645158Z digest=sha256:1c837f02cc0a24a3c3d1e11f11118f3b17381699bc41d0d76f28a16a53aade05

Observation 8e3347ca-54f3-4957-8d95-4b6af5c0d62e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Measuring Massive Multitask Language Understanding

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.647318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.647318Z digest=sha256:b719f5491c09fd45408592944d743296d2888fc44f19f95f2dd6f3d7814b8948

Observation 00071c23-c419-4149-b7ba-be718d7fb249 · outbound

This paper cites Qlora: Efficient finetuning of quan- tized llms.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qlora: Efficient finetuning of quan- tized llms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.961340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.649851Z digest=sha256:407ffa5d90f8dd86bb5a3c11e4253b6fb1a09c7548451995d9f302475c58e6f0

Observation eb141789-1fd9-4c8f-895b-2aebe0be15ac · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.651866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.651866Z digest=sha256:67a161cd18d05c149672ca7b740154391ede27b034f0bc979cc3f890ce13eeec

Observation ccb460b1-b160-421d-a89a-934d83896715 · outbound

This paper cites Mixtral of Experts.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Mixtral of Experts

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.654525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.654525Z digest=sha256:13e53cc56601161bb76e1d6dcc870b435244242606a15b5727b507a5cbf254b7

Observation 5bd9bd37-2bfb-4474-83a5-fff8d3c9e48c · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Zephyr: Direct Distillation of LM Alignment

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.657148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.657148Z digest=sha256:0a4181e787edbe2571e17c37842b014b4581360310aa943384aff6ca4b02b9a3

Observation faabd1bf-74b9-4db4-8168-81affad895e7 · outbound

This paper cites Qwen2.5 Technical Report.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Qwen2.5 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.660075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.660075Z digest=sha256:38c14f1acd8e61a2c69b858fc9e2261b4219831b32c3a2471d722e7b0510857c

Observation 5a5aa9ea-1b8c-4af4-8fb0-b5de3faa4aee · outbound

This paper cites Hashimoto.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Hashimoto

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.662669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.662669Z digest=sha256:896c143fc2d799fc7089205955bfb23b2cceacd2662f0aab78471b4628c73985

Observation 8efe3443-9481-48b6-a96d-ff673f06b509 · outbound

This paper cites Language models are few-shot learn- ers.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Language models are few-shot learn- ers

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.665104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.665104Z digest=sha256:0e59e46b1192ca28addebe337b227cf1e1efed57c119b0bf65be1264cb26fbad

Observation 4a7600c1-c273-4d04-9415-e4c8c40dab33 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.667643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.667643Z digest=sha256:2854e9b1261e62a8a360aa96033782f348c05a9fc3db9a9c6a414b369b1af85f

Observation 553ca044-2223-4d81-820c-df4a5472aaea · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.670171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.670171Z digest=sha256:4cf9d157993c955060eab51ff56066dfbd903c08bede427ae726a2be6258097c

Observation 8404c6c5-62b5-44a5-b9bd-9a2607544e79 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.672901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.672901Z digest=sha256:488657a0139d8aa056a6ae90839364d50a8adb563913f465a67bbe5e51b18c09

Observation 5f45c01d-dfe1-44ed-96e7-a634fb74bb6f · outbound

This paper cites Stanford alpaca: An instruction- following llama model, 2023.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Stanford alpaca: An instruction- following llama model, 2023

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:37:39.942588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.675692Z digest=sha256:a745d1126e1edbd168fa34cc7ab0a43f41d1d5ee24eed27124dd2aaa24fbd21d

Observation fa618968-8898-4dca-a5ae-b76e4f6f171f · outbound

This paper cites an unresolved cited work.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:37:40.201564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-09T17:37:39.500970Z digest=sha256:e0a8c6fb90efe574f47cbd4e2672803eeec6c69819e779ef7bd15e3ea2556c96

Pith citing papers

Observation 80a66c5d-3841-47f2-9391-cd8c2d36cf8c · inbound

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models cites this paper.

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:19.630547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:19.630547Z digest=sha256:ba59875b0907fc3f99d18c69281bd54f66272e72933d39fce53727896c717325

Observation 4c9d6148-5e6c-4d0a-b445-f368849463d7 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:10:21.874478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-04T23:10:21.257150Z digest=sha256:5af60738d5a7246799aba76c6b96ef911d11a44d78a89655810e95c8381cf119