Pith. sign in

Paper Citation Record · LEDGER

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations

As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.06043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06043 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:18:47.341663Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:08:04.673125Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:08:06.388576Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20ab7275-4286-4557-9065-dbc99fd0f870 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.488602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.488602Z digest=sha256:aca1b147167b54be08c3cb488294c35765d63e25e04e4cdb6da309f58b9b893f

Observation dca1a89b-0b75-4bdf-9f69-555eb9834939 · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.524938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.524938Z digest=sha256:4c23028bed305bc2ae1247ab05c73e171e3edb9eb22a5a605367b96a58f29d48

Observation 05791a07-f465-41f7-bb98-f0b59ef38b68 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.645786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.645786Z digest=sha256:1f796f89f015ed18f8fcd6c6d00ee07e925f737827d47dbdd305ecf3360b3d25

Observation de9f5339-3746-4348-96d9-378f28bdd6ab · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.749367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.749367Z digest=sha256:270e89a2ab43f103bf4714d3dfaeda38e84d5a60f6bfef31d4371b0d9adb0af4

Observation f211f1dd-e80a-4b48-9029-1116493f6a2a · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.849506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.849506Z digest=sha256:467f1c4a9513be8593fea6f44bff1bb1aa13f1470040165110859fe4517870c8

Observation 155f482f-eb0c-4b09-930b-c5236f1f9e56 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:42.941390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:42.941390Z digest=sha256:7d816580776aa8b40d42f9f734394167484d2945ddcf1237ac09f2f1d7220abf

Observation c9a0a10f-545e-408e-a876-ba2746afa4df · outbound

This paper cites Recent Advances in Attack and Defense Approaches of Large Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Recent Advances in Attack and Defense Approaches of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.023171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.023171Z digest=sha256:c15984251afd4f57d92957ff450b827cb99ae65bc910bf8e53427280d778db81

Observation c2213e70-c872-4395-afc6-7cd63728a498 · outbound

This paper cites DeepSeek-V3 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.106007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.106007Z digest=sha256:03aa9f1d406bf05d1e2a88f22b3e3950bfc11d6abdb8a396ae06d0090375d5aa

Observation 66a459e4-4b2b-4c9f-9f7c-cc657cb21a57 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.155123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.155123Z digest=sha256:35543f056aa081f20ae8c051351ec0224e6af55548fa5f185b61e692da411ccc

Observation 90cbfe8c-bc9c-412d-934d-4bf66a9f80dd · outbound

This paper cites The Llama 3 Herd of Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:43.324924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:43.324924Z digest=sha256:be7349f9d1783f21faa44ee889355e00c6a165da7e017a045038263c3bb3451b

Observation 2a34fef4-bc96-4448-b0cc-ee6154146cec · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.761528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.494982Z digest=sha256:1ca45a24dd02391b1cf3bf985431f33d7d78e35bccf91e033c99870d8fccecde

Observation 14825689-1475-4b1b-92ab-944630b02ac0 · outbound

This paper cites Cai, James Wexler, Fernanda B.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Cai, James Wexler, Fernanda B

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:49.602687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.629083Z digest=sha256:7ae987757e8cc673d6fdba5dfdbe1d0e3ec7ee8705efdc07385546b67f0de92c

Observation 2baaf5bd-1d05-4f98-852d-db0b98c6f7f7 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.424805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.749206Z digest=sha256:1bdfad0e00f400af0a3b15a4ac2114a07db1a60860f2079977c32c89b25f5194

Observation ee87e432-a3ea-48b7-8126-f346b8db9b00 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.237439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.839587Z digest=sha256:b527221a5688fd90545ba553b6e570eb877feda51084965cc9c241808398ac33

Observation f4e61275-a567-406e-9d36-772ed10f644d · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:49.039356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:43.957717Z digest=sha256:351490cbc5998aea1f9c6a8ababc0bd6866ebcef23359c352fd7624129a77170

Observation ffdae9da-ff4b-4453-b3b6-53c2e2961b41 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.110225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.110225Z digest=sha256:8a788899a073ecfefa273334a17aaa3deb803767db7dcca614e4779bf2525a7b

Observation 60fcaa8e-59eb-4e78-9b0e-88e058a1dce9 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 17

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T19:18:48.293208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:44.315764Z digest=sha256:dda6a7e003549c81767d9ffb692bc5082909661a2bfbd0c7defe782e1c6c4e40

Observation 92ec32f8-01f8-42ac-9f86-3b9e0d7f35e5 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.434891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.434891Z digest=sha256:096ce2616e9cc280d95d162b2dd4c9b61adfeb6ed93511f756006695790331ac

Observation e75817f1-eb76-4a4d-ad33-e6183d1af717 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:48.877206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:44.557865Z digest=sha256:788be4debfcd0ebee37c82a823cd1c78f474633e817a0db7b58d8979761419c9

Observation aa26c873-6fc4-4540-aa2b-924dd1d18413 · outbound

This paper cites GPT-4 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations GPT-4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.729046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.729046Z digest=sha256:55e01b28dc2b063059321e2d7b10101450882c51458975935d50882fd4395639

Observation a67b6306-0fed-4213-9137-12a769f00d0f · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.823274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.823274Z digest=sha256:103485a56458dba01f1cdb41c3503d600d2f9f140412062d7298ed8dfe2ea0f2

Observation 79a238d2-a2d6-4640-99b8-806ff726846c · outbound

This paper cites Qwen2.5 Technical Report.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:44.931351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:44.931351Z digest=sha256:6085ee6fd6d59f5159a40136ea818d3388c04259c3e6e185a37db084fb69c2a4

Observation 39ceda5e-54fd-49f1-883c-711e3adea093 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.033818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.033818Z digest=sha256:f70695ad40804839ccc2141d519d2e65732e761dc0652d931f867616e4972da1

Observation f26780fa-9390-45b8-bf79-0ac77c907ffc · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.148774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.148774Z digest=sha256:7da044dc98fc578f6635d88c3079f7c18edc359f23a38d70259a5595b0b8ead2

Observation b578c107-79d4-4f86-b42f-35891a32855f · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.243985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.243985Z digest=sha256:a5c8e3fd00f1533da48b541e13a99dee8534fba8d160625032b0c6adbf057ece

Observation b1715f98-6308-4bfa-a4e3-663672797a83 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.379143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.379143Z digest=sha256:589265bc6958273b51c87c3abbdcaf2dc7edcc8bf21841570791b4f147236175

Observation 5284303c-4ce3-43b0-a30c-41b934161202 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations A StrongREJECT for Empty Jailbreaks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.500292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.500292Z digest=sha256:6222a3d0bca3c165a99042805c7aa27361ca000bc1849c46cfefe0e20d072881

Observation 8d8f141d-59c6-49ae-ba2b-348faa9e9ffc · outbound

This paper cites Hashimoto.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Hashimoto

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.634496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.634496Z digest=sha256:bcac42055373644418857cd5a56f6dcba9be331c941443e44cd06c11647d8db8

Observation 758c82ee-0cc8-4b74-a19b-d0aaa8d1712a · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.728426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.728426Z digest=sha256:b4785bf67db810ff0ab4683f7f265a8fd79bf48791586662164629d4d218455d

Observation bce1bd3a-1c87-4860-93ae-a6b77a32d2df · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbroken: How Does LLM Safety Training Fail?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.849258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.849258Z digest=sha256:4b5519322f81695ef49329c8e883e538cd177a23de765e5bb3d7c51484fdf400

Observation 0a191a1b-a55e-4e05-8ab0-ed4870e40c1e · outbound

This paper cites Dai, and Quoc V Le.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Dai, and Quoc V Le

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:45.952671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:45.952671Z digest=sha256:c1d767bbe78f9380ba0a972b359c92f241ae5a4b232c40c3e1ddad642cebb0ef

Observation 8f7b9bf5-41b1-4bcc-b490-cae65e440cd3 · outbound

This paper cites Ethical and social risks of harm from Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Ethical and social risks of harm from Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.047757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.047757Z digest=sha256:11caf0f42591474b853b3e59b7e49e7d08d90b7d9d3006204fa42aa8a5003b20

Observation dd2c71f7-cde5-4944-bc27-36a99e3ed40e · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.137678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.137678Z digest=sha256:8b9bf41e2de330bb880565a9c53e25ab788bace5c36aa75b94c282f2fbd9bd8c

Observation 65c38dfd-c68d-406b-834d-8f98bd11a601 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:48.613526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:18:46.223411Z digest=sha256:d5d4793e267ae53399939db2721818c98e35444e7778dc5fc0c046991d337551

Observation 2bca0902-16d6-4263-b795-1ee6f7ec13bb · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.295508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.295508Z digest=sha256:28d677a64042c442a7ef62e16e91d579eb5bb5710a9eb8eada87e9f2d2e5a6fb

Observation 84632e61-765a-4cf8-b5b5-9911eab0c50f · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.428000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.428000Z digest=sha256:3e5ba6dc3eb3218165acc6195a436467888dd473cf216f53a6d1a152004cb372

Observation 70640b84-e213-4601-b2c0-aa0b955aa74c · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Low-Resource Languages Jailbreak GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.549467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.549467Z digest=sha256:723c8a5a2672141139b6220d1c67ef40f83a497c9f82643ad2c1bc7c754757aa

Observation 87b57c8e-c12e-49b7-a269-950a04fd0d57 · outbound

This paper cites How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.662354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.662354Z digest=sha256:f675922584553d4a73b23f4a6ae859df42c84e6aff822e23073c477b43893113

Observation 36385948-cc43-4a26-8b82-35896d8f38de · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.786950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.786950Z digest=sha256:5c722a4b826fd23c28835754221997d9f831319c4065dc0ff5bc7dbd53b8e46b

Observation 66a5804c-728e-4412-a414-3aaff9b71d58 · outbound

This paper cites an unresolved cited work.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.864156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.864156Z digest=sha256:276b0faac50bb78db3c77f5a0fbde23ee3eeadd1a608ef39d9ed95c97cc5ab60

Observation 0f3429be-0dbf-4ac4-8253-5162e656584a · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Representation Engineering: A Top-Down Approach to AI Transparency

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:46.970519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:46.970519Z digest=sha256:cb74c76b80c9d5a455b191b752fb1bdd663fc4562a8e6b6401988d1135790922

Observation 822a0e55-2997-4ae1-8622-df97a4992f59 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.094379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.094379Z digest=sha256:ab3699b499ca8ca7e336798eea94edfddd08f87b7201fbe380ab8dcf773a1cdf

Observation 86dd1667-75f1-4985-ba75-b247edf29fbc · outbound

This paper cites online" 'onlinestring :=.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations online" 'onlinestring :=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.235990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.235990Z digest=sha256:0f8158cc5547546102b3e18b527c81d57bb092fdf4dd57c5f6c8a3145bd034cc

Observation 13d529b9-ac61-4a1e-8e94-43b2e793f351 · outbound

This paper cites write newline.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations write newline

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.341663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.341663Z digest=sha256:f629ce794b6dfd8ee167e41f1c108d6878486a9af127a66d5370f1e0bdaaf9d9

Pith citing papers

Observation 7c1330a7-e8ce-46f9-b491-83cbe88844a7 · inbound

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs cites this paper.

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:08:06.426230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T21:08:04.673125Z digest=sha256:7584fcd1438404d581762ff08537e45d5825867475e582216a8ee4a6d5353ab3