Pith. sign in

Paper Citation Record · LEDGER

Linearly Decoding Refused Knowledge in Aligned Language Models

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2507.00239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00239 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:35.040995Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T05:52:53.880521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:53:04.283223Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact5
  • verified fuzzy22
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5a2a791-7579-49c3-88f6-b2715a52fc5a · outbound

This paper cites Yi-6b-chat.

Linearly Decoding Refused Knowledge in Aligned Language Models Yi-6b-chat

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.934291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.845006Z digest=sha256:756523620f3f067f9a23c07e33508ba110e999a9843c792753dc3628bccc0ec6

Observation 60715637-4fce-4c16-a3de-03f09fe8e2c7 · outbound

This paper cites Fine-grained analysis of sentence embeddings using auxiliary prediction tasks.

Linearly Decoding Refused Knowledge in Aligned Language Models Fine-grained analysis of sentence embeddings using auxiliary prediction tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.925686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.848651Z digest=sha256:baef151f944c5aeb1e341d4138ed0ebc2e8e411a7774a8be13f04305a0e66502

Observation 8969fc5d-c558-4fd1-8a32-157c314a5060 · outbound

This paper cites Understanding intermediate layers using linear classifier probes, 2017.

Linearly Decoding Refused Knowledge in Aligned Language Models Understanding intermediate layers using linear classifier probes, 2017

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.851907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.851907Z digest=sha256:bfd27f273c5d4ac1e864dbc390c1f78e6110b50ad85d9bb6913dfb4380955be0

Observation 69d7fd01-27f0-4eb8-9c1e-9955d628c309 · outbound

This paper cites Bowman, Ethan Perez, Roger Baker Grosse, and David Duve- naud.

Linearly Decoding Refused Knowledge in Aligned Language Models Bowman, Ethan Perez, Roger Baker Grosse, and David Duve- naud

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.911881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.854780Z digest=sha256:68f362538f4dd75950e6d96604a22b62cccc16c6b33b1291afd27d6ad8584835

Observation 37a6811c-38e0-4a5e-9068-f4c2007e0ea5 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Linearly Decoding Refused Knowledge in Aligned Language Models Refusal in language models is mediated by a single direction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.904069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.857729Z digest=sha256:a2777296ded12da59283ff3cc67af487c9c0805176f7100bf330f68f1916aedb

Observation 9c846136-d36f-414a-a8f2-ef648efbce58 · outbound

This paper cites Language models can predict their own behavior.

Linearly Decoding Refused Knowledge in Aligned Language Models Language models can predict their own behavior

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.860566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.860566Z digest=sha256:647d55f8e622875c407d0e350293426656c0b0d2735dfc79aa85edf662bfb355

Observation 485bfb6e-dec1-4c15-9d9b-1a77e5774d35 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Linearly Decoding Refused Knowledge in Aligned Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.863565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.863565Z digest=sha256:eb0efd9c9f44ddeab79f38bde7ad07c016d9be313a32932e5c9d89bec6159075

Observation 2498cc3d-3119-430c-99e0-fa4527e14201 · outbound

This paper cites Probing classifiers: Promises, shortcomings, and advances.

Linearly Decoding Refused Knowledge in Aligned Language Models Probing classifiers: Promises, shortcomings, and advances

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.866492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.866492Z digest=sha256:f12b72df707b2ed660672f149a84ee0eb94abbc4f365ff5ee6e35140064f5078

Observation e12569d5-02ac-4d67-b948-992ae2bca332 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.

Linearly Decoding Refused Knowledge in Aligned Language Models Emergent misalignment: Narrow finetuning can produce broadly misaligned llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.869282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.869282Z digest=sha256:c597764392a237e53f323ee1f00ccb2ff7d69d09d4e353a06d1551a78d83ef00

Observation 31e68641-64d1-482b-96b7-34c604c2b306 · outbound

This paper cites Wedded to prosperity? informal influence and regional favoritism.

Linearly Decoding Refused Knowledge in Aligned Language Models Wedded to prosperity? informal influence and regional favoritism

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.895949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.871895Z digest=sha256:3ad11513f2ef6a18f709441502f023c74c1a252414fb48c8d054eda8f491a182

Observation f0016ce6-8359-43f6-85ea-e5843884496d · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.874834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.874834Z digest=sha256:2869452743629aa84ed941c604f4ca36dc2cdac668cd4720dc1d49daa98399fc

Observation c2aa32b7-6cc5-4ae0-9e33-b4778e0f59a0 · outbound

This paper cites List of countries | Britannica.

Linearly Decoding Refused Knowledge in Aligned Language Models List of countries | Britannica

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.887726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.877444Z digest=sha256:570e4848988aa59b2da7b88efc5a959485fc05e4c149aab6f7bc30a0312c2413

Observation 439e7a25-52f9-4974-a467-8996553ed81c · outbound

This paper cites From Imitation to Introspection: Probing Self-Consciousness in Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models From Imitation to Introspection: Probing Self-Consciousness in Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.503514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.880028Z digest=sha256:c4022ceec180efcbdc53506d44e7ef042a1b1ed0ea10be71e8409e20e44d3410

Observation 950eb059-ec69-4cd0-a671-ddaaba6b4324 · outbound

This paper cites Probing linguistic information for logical inference in pre-trained language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Probing linguistic information for logical inference in pre-trained language models

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T21:27:35.112492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.883871Z digest=sha256:f53de09cec071d93744cc22c9441de071aa5e89f02ce98da14f8fadbe07f3113

Observation eefa3a7d-cae6-4863-a024-f84dc69db856 · outbound

This paper cites Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks.

Linearly Decoding Refused Knowledge in Aligned Language Models Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.886612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.886612Z digest=sha256:65c0a1ab7df4b40516c8a89ed0630320adf63c3fbe1ef6ead593e241c634cb45

Observation fe531612-2443-4d1a-aa94-e4a8fd5900df · outbound

This paper cites Breaking down the defenses: A comparative survey of attacks on large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Breaking down the defenses: A comparative survey of attacks on large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.889561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.889561Z digest=sha256:bb75b16eca355f04291831b828465e182734b670e561888a7a73494b2b10f960

Observation 01705d48-dbcf-4a46-bc10-2a5b4228527b · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Linearly Decoding Refused Knowledge in Aligned Language Models JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.892626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.892626Z digest=sha256:a80cebfbe16a8a0bb1981fbdaa3d97b5881648e503b15abb522b433c12a1d0d5

Observation 526a21f6-bf23-4ce8-80df-4186450bc962 · outbound

This paper cites Scaling instruction-finetuned language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Scaling instruction-finetuned language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.895935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.895935Z digest=sha256:b92fb1556b3733970b06dcc9842ecc080750b9b043a865b043e1775422c74100

Observation f180b5d0-bf86-48b5-9d09-b5531dafd636 · outbound

This paper cites Pawan Kumar, and Adel Bibi.

Linearly Decoding Refused Knowledge in Aligned Language Models Pawan Kumar, and Adel Bibi

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.875007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.898752Z digest=sha256:ddc3b86f1efc6a20825b8ec6c1720e3d81ed5fb46e1793eada4cedf4b6ba22df

Observation b092e29e-e050-4f37-a99e-6c79a296a0d8 · outbound

This paper cites Dissecting recall of factual associations in auto-regressive language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Dissecting recall of factual associations in auto-regressive language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.901466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.901466Z digest=sha256:416fa1ca0ee9385a57e2eb668665d5b16ab312089f1069fe36aef0413f9e71a7

Observation 06d3705b-e684-468a-9d09-dfba78d41227 · outbound

This paper cites Estimating knowledge in large language models without generating a single token.

Linearly Decoding Refused Knowledge in Aligned Language Models Estimating knowledge in large language models without generating a single token

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.905010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.905010Z digest=sha256:be636eb5d94abd20ff88320b0d9d51a39f73c932b1fc5478856dd0036aceb0a4

Observation 6fae92c4-ed83-4aa5-b882-635b997b8616 · outbound

This paper cites Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.

Linearly Decoding Refused Knowledge in Aligned Language Models Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.908262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.908262Z digest=sha256:9070e16e54ed50d768907170b4627278fe2d070d73bab1d39f4077e4e03db8d0

Observation cfa06f62-615f-40ca-9ed9-45d4fdd293ab · outbound

This paper cites Language models represent space and time.

Linearly Decoding Refused Knowledge in Aligned Language Models Language models represent space and time

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.867466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.910913Z digest=sha256:56f03b4a658b7530d031b15be40e41839da6d4b9e3ee282390a4a5d0e6edc50e

Observation 646a3c4b-2bd7-4830-8711-db5f719a9414 · outbound

This paper cites The Elements of Statistical Learning: Data Mining, Inference, and Prediction, volume 2.

Linearly Decoding Refused Knowledge in Aligned Language Models The Elements of Statistical Learning: Data Mining, Inference, and Prediction, volume 2

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.859326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.913521Z digest=sha256:c003e7dd0a39f7657d78a20f828ac55c8c729aa5b3f0b3c7750cc65f7638c7bf

Observation c3ac457d-fb07-4073-bd67-6fd10c97a6b4 · outbound

This paper cites Do LLMs “know” internally when they follow instructions? In The Thirteenth International Conference on Learning Representations, 2025.

Linearly Decoding Refused Knowledge in Aligned Language Models Do LLMs “know” internally when they follow instructions? In The Thirteenth International Conference on Learning Representations, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.849401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.916396Z digest=sha256:ff2391da138d42cb062f479dc684bf2df4676a9f72301c5d89e7cdcaaa6eab33

Observation 759b8f54-cd1d-4747-854f-355b875cf830 · outbound

This paper cites Linearity of relation decoding in transformer language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Linearity of relation decoding in transformer language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.919014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.919014Z digest=sha256:d1f6a10f2e1e0dbb9f16727086be64d4616265bf1f622afe6601dd62a1b2dec5

Observation 8430ffe5-7061-4fc7-9eef-39105ff4c5cf · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.921758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.921758Z digest=sha256:197090c6d2fccd73b2d4d9e1a5018876f61895d74669d8592721eb009fb9dc71

Observation c01ad35f-d372-4f9f-b395-75534335cdf1 · outbound

This paper cites Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.924646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.924646Z digest=sha256:3e205d25cf66c9fa3afd81a37be29dc29a93712b6b3b97d1fa5597e592d303ca

Observation 3336f87e-aea9-41c8-a87f-3a981ccda430 · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:27:35.836329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.927770Z digest=sha256:884f40b06fa73507ea4be3ae352347771c9e21bc7ac8e736bbd1d0f6a1984ca8

Observation cfda084d-86b9-4854-91d3-8374c95541d7 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.930716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.930716Z digest=sha256:1d133eb316c7e1c83ff721f679a8e0671e8fb49cad9203f7693fa23a38b2860a

Observation 4bbaf33b-6b53-4065-8505-50fe0f78aa21 · outbound

This paper cites Alignment of Language Agents.

Linearly Decoding Refused Knowledge in Aligned Language Models Alignment of Language Agents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.933507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.933507Z digest=sha256:eaadf72e656ac52a879ffd7d6ddbdc722669e415270e847fa4624f949f60486c

Observation 21901eb9-72ba-4d6c-903f-2f333a024a6a · outbound

This paper cites Linear representations of political perspective emerge in large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Linear representations of political perspective emerge in large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.936952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.936952Z digest=sha256:76bc856f18b8e251bc9a59c9eaa0a47b418014f0f08d8a0731bad9f7eeb3ff57

Observation 7baf0ade-4903-4696-9913-52d67c38e2f1 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Linearly Decoding Refused Knowledge in Aligned Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.939664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.939664Z digest=sha256:28a35e909c4d675c6340e723792320af0b46dc57d8f2820ead877da271894b51

Observation bcc200e2-fc7b-4a7c-aa80-f5ae16cb250e · outbound

This paper cites Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.231753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.942636Z digest=sha256:5f29dda37f047b32013713c1f4d8718052ace7470ab3da9a381c520a8a17dcb1

Observation e8c6d6a0-da41-49f0-b24a-5a9ff6d632cb · outbound

This paper cites The unlocking spell on base LLMs: Rethinking alignment via in-context learning.

Linearly Decoding Refused Knowledge in Aligned Language Models The unlocking spell on base LLMs: Rethinking alignment via in-context learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.823624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.945645Z digest=sha256:d512267d8a14fdc7b58ae6f5df17fa13609a1c18aef5a59d1cc29c3d54921462

Observation 1dabd6f2-50a8-44ae-8f45-764cd3e04cc0 · outbound

This paper cites Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis.

Linearly Decoding Refused Knowledge in Aligned Language Models Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.948426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.948426Z digest=sha256:869a4b4319899f334dd48e4d34c9c870aa78208f60c7b9269b9c8ac23b06d2cc

Observation 74188ff8-dbf6-4b71-9309-9484a5e2fb27 · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Linearly Decoding Refused Knowledge in Aligned Language Models Formalizing and benchmarking prompt injection attacks and defenses

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.815303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.951313Z digest=sha256:7fae1cea469ab4b151e0e07d3b7bf7408590013dd4e4ed703bea726989469045

Observation 70fcb23a-e528-4821-9f89-888045f82745 · outbound

This paper cites Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates.

Linearly Decoding Refused Knowledge in Aligned Language Models Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.954267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.954267Z digest=sha256:1a06e50489b92935a0a1fcbc09887e9a8cd6628f865a0ea4cc0b332b3fc9f0b1

Observation 86438a36-a411-4441-a471-54d8bed4c87d · outbound

This paper cites The geometry of truth: Emergent linear structure in large lan- guage model representations of true/false datasets.

Linearly Decoding Refused Knowledge in Aligned Language Models The geometry of truth: Emergent linear structure in large lan- guage model representations of true/false datasets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.806655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.957079Z digest=sha256:dc24f3a96ce7f1d4feed51cfef26801bfd99526269f598be399fa8eee717bda4

Observation 6c34073d-86c8-4a1e-89c7-b423059d414c · outbound

This paper cites Occupation Data - O*NET 29.2 Data Dictionary at O*NET Re- source Center.

Linearly Decoding Refused Knowledge in Aligned Language Models Occupation Data - O*NET 29.2 Data Dictionary at O*NET Re- source Center

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.789994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.962407Z digest=sha256:de6a3971e91c6de57d4528de0a67dfad108384e53d8a7d4279c61cc840f9a2d4

Observation b186f49c-7811-4216-8e8b-7da273136424 · outbound

This paper cites Training language models to follow instructions with human feedback.

Linearly Decoding Refused Knowledge in Aligned Language Models Training language models to follow instructions with human feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.781795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.965216Z digest=sha256:c336c5ccf8f09129194897b0dbb7a5c23ffe581df7ec96b20eb32f0d5ee0f8da

Observation d3d0f1e2-31dd-4717-8934-f78e7e3e7546 · outbound

This paper cites The linear representation hypothesis and the geometry of large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models The linear representation hypothesis and the geometry of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.773466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.967915Z digest=sha256:f5d2974f43f30e495e94d1d628366b930e89b40cd65edc291a522e378dbc4851

Observation 5096fc8d-d834-4727-8e4f-34e49c62aa1b · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Ignore Previous Prompt: Attack Techniques For Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.970768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.970768Z digest=sha256:7b1872db239ce59609173d0985844bc0de376138c458bdbcc1d42849283a59a8

Observation a17f31f2-5cce-4416-a8d5-5f571d90d4fd · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations , 2024.

Linearly Decoding Refused Knowledge in Aligned Language Models Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations , 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.973853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.973853Z digest=sha256:e5e56bfa3929bbd9a3aa15f88d91f52f8aeff900630fad8735cd781ac3c23105

Observation 67379375-b768-4557-a664-7adce87c13e5 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Linearly Decoding Refused Knowledge in Aligned Language Models Safety alignment should be made more than just a few tokens deep

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.977277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.977277Z digest=sha256:d34758fbeb69dc44aa06146c37221595f0272ba7e1ec124f425d2f6673bd298f

Observation fb4c71d6-3951-4f95-a5a5-ad814441f6bc · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Linearly Decoding Refused Knowledge in Aligned Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.980167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.980167Z digest=sha256:edf169369143df16a83cb65cd6a148a8ac189f64a8a9ab06357c6116bad5542c

Observation 1c158573-9d21-4878-b441-ee83b5fc65c0 · outbound

This paper cites Multi- task prompted training enables zero-shot task generalization.

Linearly Decoding Refused Knowledge in Aligned Language Models Multi- task prompted training enables zero-shot task generalization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.749217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.982721Z digest=sha256:1b3bb7c2cb6f325236d76de95aaa6cdb6a14d9750b4e562d6f1fc75c8d5a0ee9

Observation 74e33e9a-f10b-4f6d-88cb-9af4f587b0f0 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

Linearly Decoding Refused Knowledge in Aligned Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.985261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.985261Z digest=sha256:343d5d352637d097a8008558bc7f8583e32c33c8d0aef49b92490dd097eebf0d

Observation 8e459ab7-d98c-43af-9a8b-2b7768a272c4 · outbound

This paper cites do anything now.

Linearly Decoding Refused Knowledge in Aligned Language Models do anything now

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.740959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.988437Z digest=sha256:dd4759aec62dc1ed9f7f5b6adf7e05822fd6edd85f20b2da14245d6d19893f5e

Observation 8ccd69e5-4941-4fa4-ac29-f2fcc1404203 · outbound

This paper cites Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations.

Linearly Decoding Refused Knowledge in Aligned Language Models Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:27:35.197892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.991094Z digest=sha256:f67ce57c99fdb672bf2d5d8fc30c91e5f26d995d487469c6a48e43dcb3f80706

Observation 89b52122-a299-46f1-ac45-1adf19bcba0f · outbound

This paper cites Large Language Models are Inconsistent and Biased Evaluators.

Linearly Decoding Refused Knowledge in Aligned Language Models Large Language Models are Inconsistent and Biased Evaluators

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.994224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.994224Z digest=sha256:34b8ef31afbc102d0726dc13bb89bd08bb6c4a5710e3a0078125633f8f2fe064

Observation a88fe132-d912-464b-adb0-575e52fb0356 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Linearly Decoding Refused Knowledge in Aligned Language Models Gemma 2: Improving Open Language Models at a Practical Size

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.997511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.997511Z digest=sha256:27d0149dff24663fe7ba27ffe4a4a76d370e30dd39280f40a5536ba42a70140d

Observation 358b4c94-e01b-4aea-9bf5-e6661b3756af · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

Linearly Decoding Refused Knowledge in Aligned Language Models What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.000447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.000447Z digest=sha256:e5b599e46440b831346e6f8f83f8e286d3d20e7912c015a4d6a9ca5833a8367a

Observation d4edba58-579f-41fa-95c9-c325e80519d1 · outbound

This paper cites Attention is all you need.

Linearly Decoding Refused Knowledge in Aligned Language Models Attention is all you need

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.732940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:35.003388Z digest=sha256:122dd452767fffebeb3f11758fb99f59e1b8314117e838a13673297f1c4eded4

Observation f93403d1-7fd8-425f-b47d-251dcb60d155 · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models.

Linearly Decoding Refused Knowledge in Aligned Language Models White-box multimodal jailbreaks against large vision-language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.724603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:35.005888Z digest=sha256:f5e2d201047dc8e8bab9e19f0a922628930dada80774e755530ded4210bb4e33

Observation 8a797ac8-2e88-4f95-8733-6c1002f776bd · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.715626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:35.008684Z digest=sha256:f74d70a5c73a54b293bd262fd593e59685fc9de8d48e5435308b8d69a11154fa

Observation e52b27ab-0ec8-40c8-8109-dde23d632199 · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

Linearly Decoding Refused Knowledge in Aligned Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.011420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.011420Z digest=sha256:952e2992d408b6c874e2282dfa39b00d35912aae8a691804f3a1081fab558c3e

Observation c8564b34-057e-407c-bf81-6ed1135155b4 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.014424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.014424Z digest=sha256:e8bdb7cb4acd8b932014b6929c7ecd8954a8202347dfdd9404876566b8520fb9

Observation e8c3f6af-81dd-40ef-ab5a-21b0092acb6e · outbound

This paper cites Efficient streaming language models with attention sinks.

Linearly Decoding Refused Knowledge in Aligned Language Models Efficient streaming language models with attention sinks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.017387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.017387Z digest=sha256:8e2e67a23d820281dc51bffae1ffa7d9b89a4db653ade94b77eda0ad2e7ed97d

Observation f17e0bc1-e48f-4aac-ab21-9c93e8167193 · outbound

This paper cites Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility.

Linearly Decoding Refused Knowledge in Aligned Language Models Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.020100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.020100Z digest=sha256:02427fda59c3566f61aca0da1629d873e38156f56a268d3a1f206aa6a53d6215

Observation a33b0e54-ad19-4232-89ae-fc87d627e05a · outbound

This paper cites On the vulnerability of safety alignment in open-access LLMs.

Linearly Decoding Refused Knowledge in Aligned Language Models On the vulnerability of safety alignment in open-access LLMs

Reference 61

Resolution
verified exact
doi, observed 2026-08-06T21:27:35.073634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:35.022967Z digest=sha256:0d5f3dc21762f5fe436e09f7e482fcaf578bd1a0f586ab2c83b632c2465854c1

Observation 549a16a1-0b82-45bf-83b0-0e8a7a9e5e49 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.025865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.025865Z digest=sha256:2463323aa9a1fa028e394173e4710d0bd398275383d20048a978c2b9a8b58db9

Observation 9836f397-8f1c-4026-b996-27b1613c361a · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Linearly Decoding Refused Knowledge in Aligned Language Models Yi: Open Foundation Models by 01.AI

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.028823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.028823Z digest=sha256:8d316da5386039a994f014ba2fffbd284892680ecbc15f65838784c41451b3d8

Observation 5857d7c6-e88b-4079-bdda-232f7a5b79d9 · outbound

This paper cites Don’t listen to me: understanding and exploring jailbreak prompts of large language models.

Linearly Decoding Refused Knowledge in Aligned Language Models Don’t listen to me: understanding and exploring jailbreak prompts of large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:27:35.702045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:35.031749Z digest=sha256:4b36b9bee6acc6471239c0d2d13e2508c949c93cad956e6a25c524d67b4e32e5

Observation bf6791de-0c35-48c4-9af0-7d7f436b471c · outbound

This paper cites Removing RLHF protections in GPT-4 via fine-tuning.

Linearly Decoding Refused Knowledge in Aligned Language Models Removing RLHF protections in GPT-4 via fine-tuning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.034422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.034422Z digest=sha256:f69dde3bdc6db9bd493614a314865d87493b12259f245c7854d4dca9b44c3f1b

Observation dab3bda0-c3ff-45ee-b08c-c3cb7706f90b · outbound

This paper cites Lima: Less is more for alignment.

Linearly Decoding Refused Knowledge in Aligned Language Models Lima: Less is more for alignment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.037339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.037339Z digest=sha256:a7b50807b056b056dbaa845e17420c38dd4e44167cf68c0f00e032579d747b4e

Observation e3b2af77-ec94-45f6-9f49-31afebf2d95b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Linearly Decoding Refused Knowledge in Aligned Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-06T21:27:35.040995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.040995Z digest=sha256:5add91f19f2cd6e0ee458968e91e0f4148be490659bd8bfd13c9c692f5e86f12

Observation 7b418754-5462-49f6-8997-fc1ee1e73aef · outbound

This paper cites an unresolved cited work.

Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:27:35.798529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:27:34.959696Z digest=sha256:827eadaba7fd1013a7ff17059e037197145ef4edefeb3df09a3867672fce5e30

Pith citing papers

Observation dbcb59b9-70ef-4595-91e4-d4fff01c1fd4 · inbound

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models cites this paper.

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models Linearly Decoding Refused Knowledge in Aligned Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.285500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:52:53.880521Z digest=sha256:a80c2e600090fcdca642ab6aaf82e5fda33428c32611139b15630f28af8d0606