Pith. sign in

Paper Citation Record · LEDGER

Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2402.09283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09283 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:13:57.190980Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T20:58:25.807935Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 58198bde-1ee4-4db5-8467-a76cbf49d036 · inbound

A Survey on the Memory Mechanism of Large Language Model based Agents cites this paper.

A Survey on the Memory Mechanism of Large Language Model based Agents Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:21:39.904669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T07:21:39.440092Z digest=sha256:867abaad1c3bfffcc06c62e61ab1dc121e44aa53f4f362941f1bc3ac317bc7ff

Observation 4ff44bce-84ff-434f-934f-191c6a17673d · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:25.814451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:36302bdd1f74b21f7072f7f967840384ad860d921810bd6f33ce9e9c53bce0a8

Observation 2b4d91b7-1791-4e0c-8b0f-a307743f6348 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.190980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.190980Z digest=sha256:b2165b274fb0aa5f590c8f0acf2b9646d58fff57edf9f6b66c00b465dfd93722

Observation 96721cf1-8b99-466c-9273-56eeb7809bef · inbound

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI cites this paper.

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:07.138295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:07.138295Z digest=sha256:cd2b13304139033e3acfc66114c9b1053ee2b06ba524f26806ae203dcf0d2775

Observation 11287278-d19c-42aa-bd9b-95cd6ddf99d9 · inbound

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models cites this paper.

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:38.782601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:38.782601Z digest=sha256:d6ca81eda9c9e1746a38da07ad1ecf6196614cc3985154dc473eb2dbafa70f3b

Observation 7a878be1-1938-426c-8f49-4cc7ed5d7ebb · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.755719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.755719Z digest=sha256:624ce108e088638064a67f20f2780b4674982ec4579d61bf4845ae20ee3e3f91

Observation 92d87873-ae93-4ffb-9f2c-8da28b49f108 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.719150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.719150Z digest=sha256:47db1fbbdc48a1b72d985937b6bf546af01fceb75249959148d31a458457333b

Observation 6326dc9e-6c85-48be-9f1c-f5a4f72e2522 · inbound

Automated Capability Discovery via Foundation Model Self-Exploration cites this paper.

Automated Capability Discovery via Foundation Model Self-Exploration Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:55.180732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:55.180732Z digest=sha256:ab16ead39e1fd7c09a44a5954fcbde9ff114b6d3da21358cb65b3462dd80d986

Observation 5ca62124-816a-4f1f-9737-e4b7ec72e693 · inbound

Should LLM Safety Be More Than Refusing Harmful Instructions? cites this paper.

Should LLM Safety Be More Than Refusing Harmful Instructions? Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:47.195714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:27:47.195714Z digest=sha256:31ca1b0e8476286fff856971eacdbcc91c4e1e027ff7704995a539bbb5c0e560

Observation 3c486612-93d0-40dc-85e4-27353eafd775 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.628023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.628023Z digest=sha256:3397a92b58ecc63e3afe8eff9f3d865eef6c6e35a74fd71cf7b804569daab9bc

Observation 1786d7bb-51c9-4630-aa93-52ef0bd1c9b1 · inbound

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law cites this paper.

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:19.061499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:19.061499Z digest=sha256:8444b92fffbb8224dc7f600f14c9b1c584127dd4ba5e7fc15aac979c60b6cba9

Observation 30cfdb8e-44c4-4999-8133-323b19a8d475 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:45.301706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:45.301706Z digest=sha256:cbeaa36bedc792760ced60a001b37c419e1197d3056b8f9760c7504d3f2f9e2e

Observation ddf34a5f-a9c9-4178-a38c-74dfd1bdc6f8 · inbound

An Empirical Study of Vulnerable Package Dependencies in LLM Repositories cites this paper.

An Empirical Study of Vulnerable Package Dependencies in LLM Repositories Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:34.913851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:34.913851Z digest=sha256:c0c4bae1370b9a5b355929e25d7ab97ba54eba90753e98dc72169ba1a297f50b

Observation 403519d2-2a8b-4bc2-8ae3-e9693caa9657 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.432374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.432374Z digest=sha256:aa438331feb20a69faaf84d62360a5ae2bfd22e4b9e4e4659bfcdf1edf6451d7

Observation f00d6933-57f9-4176-918c-24f41037209a · inbound

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models cites this paper.

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:55:56.468210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:55:56.468210Z digest=sha256:8d8c8bf63a22d8e1061705b4e251cb5b50e893ce8c12a64909248d7b02f1e90d

Observation 089a587c-4ee2-47b9-9733-67c3c23f0217 · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.363837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:8557a79af33b35aa33ba6042598b724ef1dca394b5d9f7cdbe9577c6e6ad96ac

Observation c0c202ab-1e91-4d16-9f97-028daa4a0a44 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.943495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:9cdb6cf409e12044e5cf7abf74263580ff868b860879e7367c187bad1f704e16

Observation f44e2550-66fe-4be8-bce1-a6ddf263678f · inbound

Rethinking Fraud Safety Evaluation: Multi-Round Attacks Reveal Safety-Utility Tradeoffs in Graph-Context LLM Defenders cites this paper.

Rethinking Fraud Safety Evaluation: Multi-Round Attacks Reveal Safety-Utility Tradeoffs in Graph-Context LLM Defenders Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:33:57.779596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T04:33:04.629491Z digest=sha256:f7a77c0521bfac98171293d103da41f3a966d3053806a795213258e245b6ce44