Pith. sign in

Paper Citation Record · LEDGER

How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.05644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05644 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:15:24.614398Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T07:46:56.508546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7870e03f-04b4-4169-a6d9-619b08e956b8 · inbound

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance cites this paper.

GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:24.614398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:15:24.614398Z digest=sha256:e9d4fc86a587480e929e0263cef9d9ef686ad6d092b4be262af8e297a92a31e1

Observation 429d95b4-f72c-454b-a70f-8be9adf9e52c · inbound

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models cites this paper.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:33.346536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:33.346536Z digest=sha256:6956e666c623054041260fb34bdba300e8ba891398b3929d97c840f46c9badf2

Observation 51d00b66-1e49-4927-b73b-35dfe637e7e1 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:53.895617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:53.895617Z digest=sha256:50971e24a9aa2d4d686cda8d3cbdf3b76705c7c921c0cae127366461339ccf0f

Observation a6e098f9-c354-479b-b8f2-9964ef0d3085 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.500858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.500858Z digest=sha256:89044c5aa8461aaf55ee9740b21f70adcfbcc1c113662c8024fcae0d64ac1d22

Observation 7714dea5-9a52-4358-9262-6a615eea9535 · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:04.470748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:170cd2a053ba258cd7d4d86fe8c499eb5da35d4e0e4e95b235055853cbaad807

Observation 3eaf97ba-80cc-4432-8313-40a0fbf8663d · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:49.276055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:bcee07b51c7e3b764113358bdfff8e80bf1555faa3b2ba8b3aee56b6dc770e76

Observation 6df54325-e926-4bbe-8a01-62e1217e5e4d · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.112172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:3c6a615f20c269dbc29501686c46e62fc3b7d086908637e6d8987ccefdf1641b

Observation 38a18b8f-01ee-495f-956b-76c41180bbf5 · inbound

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling cites this paper.

SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:59:33.720230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:23:06.382761Z digest=sha256:2027d7091945bd85349bb90f460cca6e9a6f2d50369f7a856ff937c9e59c8927

Observation 3fbb7da7-e5fb-447a-8874-2ff1e1274ff6 · inbound

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice cites this paper.

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T07:46:56.509857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T07:41:35.741974Z digest=sha256:755c758ce2705591bd2534f27041c438b543619e0a46b7e0c0d065b9945517b8

Observation be95769d-5559-4498-b31b-3c400c93723e · inbound

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice cites this paper.

OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:06.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:06.547479Z digest=sha256:5012cec46b6574fea51a2c6ec17cb93dbb3d312d0cacd4f1e5c5a3a8aa252931

Observation bfd23706-d0de-4641-a148-94c13ce1ec17 · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:30.920135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:30.920135Z digest=sha256:26e1c67b2f2e9ebf64afd462ed993b80f809ef5ab63db57abb1b3d2fabe0efa9