Pith. sign in

Paper Citation Record · LEDGER

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2501.10639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10639 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:08:01.799775Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T15:26:23.290009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T15:27:20.106151Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0dddf4d-c200-4382-af12-531fb113d44c · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.566687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.566687Z digest=sha256:656d6d9d5e8917985cd1563c0fcb44fef36ccd85dbf2a6ba6f4c954ea9132514

Observation 5189c3aa-57f5-4e1d-ad00-3031b6a7f22f · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.571828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.571828Z digest=sha256:b7b23d7dc05728c8584d117421b5a310bccbb2a81c35b4d92291bcae7cadeacc

Observation 9888649f-64a0-4d39-860d-ab2b2a2291d8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.578805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.578805Z digest=sha256:17ed5147c21e22fad778e77febc6c2da1c53a45c2901414e16460329a058c0e5

Observation b189b13d-7f8c-41dd-9fd9-c0b334bbead5 · outbound

This paper cites , author Ghosh, S.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ghosh, S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.565008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.583668Z digest=sha256:e38570b3878c327cb0479a0d4df0ce58146e2adeef74d66a3b6c80090c27cb5a

Observation c67ebe50-2a32-408a-8f48-c3504b14a518 · outbound

This paper cites , author Ye, H.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ye, H

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.551904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.588539Z digest=sha256:a0d2f64a03cd2ee7560c6888bcca4ec3e3310ed09634575a22a26f8a61b8d556

Observation c7cf5986-6253-44bc-b21d-cd996d3ce9f1 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.593502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.593502Z digest=sha256:f16847c900e2929683977c35cd55c3831a254367c0238608226aa7289fd95c52

Observation d2ac95ef-8939-4c86-9984-a03b39db9a6f · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.599083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.599083Z digest=sha256:cafc32469c02aecb38fbdb594b274d3955368bf4b8a48143e9c9424cea40051f

Observation 1d04c9aa-2ef6-4740-8ab8-0f33ee4c147d · outbound

This paper cites , author Robey, A.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Robey, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.538492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.604071Z digest=sha256:b146887061213966d017417a8cfead8b4c397c40920814d70ff3cfbd522e2add

Observation c530ab1e-d492-41e7-8627-266a1c0dfe3b · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.608940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.608940Z digest=sha256:62d0437e60335ef1c4ddb142199d5dece77778968c7c62e431447fbc6f454290

Observation cd124da8-9e9e-4263-b6e3-6b983ac61d1a · outbound

This paper cites , author Ruoss, A.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ruoss, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.524902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.614032Z digest=sha256:31a5d86ecf74b75996fe6d7e0db714112cdd3e27da7612e660075525977c55e7

Observation bdfe378a-b3be-4cf4-9931-f706dc698eb1 · outbound

This paper cites , author Chen, Y.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Chen, Y

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.511111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.618703Z digest=sha256:30ebfac93a556cda21ed6e017f96b485ff8e54b4e12addf247bc34f5c9806660

Observation 97378996-f677-4d99-ad7e-a4440b3760e1 · outbound

This paper cites MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.623284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.623284Z digest=sha256:2db86c268621fb992be64385e54331de160816e64b438b0d4e8b4289729755e9

Observation 94febe71-c2b2-4e3a-8d95-b055e2ad8e64 · outbound

This paper cites PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.628010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.628010Z digest=sha256:0b6974e65b59b666d2afa15a6b964b13eb2fb4569354317b30a1b0c383961575

Observation dcdc1a55-96d3-4e61-bdcd-9bfbe6f56725 · outbound

This paper cites , author Yu, F.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Yu, F

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.496311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.632772Z digest=sha256:80d2495ffdd755054690a32a872fed361a4f1b57a67cdc8907165ee8fa0bc072

Observation 8d403bd6-9448-4d89-8a6a-d07e32812eae · outbound

This paper cites , author Burns, C.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Burns, C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.482280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.637090Z digest=sha256:ff7812e5d23c8aa0b3902d5e252918214ae476178cb209c7fee6476782771e5e

Observation fe449179-5845-4b2c-821b-386d09820afe · outbound

This paper cites Inspecting and Editing Knowledge Representations in Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Inspecting and Editing Knowledge Representations in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.641136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.641136Z digest=sha256:36b51a8dfa2e51dbabaf750346b9b6eaac2986845a18c41171a3fda8af055c4d

Observation 9c367882-5566-4d4f-baab-1155ff1a9ef4 · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.645564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.645564Z digest=sha256:95329b7367dc7a91c6c902fa3262ae6dd91ec3e63976768b012bb84230847b79

Observation f8e218e1-6d2d-4e11-b27d-620525c5b7c6 · outbound

This paper cites , author Choi, E.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Choi, E

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.467635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.649858Z digest=sha256:320d4207462979dafecf76100e8e951b281722f00ad4a691bc89434a847e7a2d

Observation d697efc6-c479-4a4c-b0ff-2a0700cfffb6 · outbound

This paper cites , author Li, X.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Li, X

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.454451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.654313Z digest=sha256:2dbe8526ea4acb2bf2c2934f31c680c6756aabe25bf10f4a4a1aeff1b81e5d0c

Observation a5e4a53d-703d-4811-b94d-203576856598 · outbound

This paper cites , author G \"u rel, N.M.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author G \"u rel, N.M

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.441977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.658208Z digest=sha256:4372628a1690b4acd590196ede445d9cf3bd7bff13f82cc1069c139867a5c704

Observation a2e3eaa3-3127-4a17-9a5e-66867126df9b · outbound

This paper cites , author Al-Rfou, R.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Al-Rfou, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.426978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.662138Z digest=sha256:8a5dc82d5c6427e72a141792df2ade56a45544a3b359861c70b830dfd4973bd4

Observation 6be3978e-c578-4853-b6b2-d25fb739904e · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.666387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.666387Z digest=sha256:323fce3343a9f0b8e1b2476709da0ceaa84258cbbed37542623c85661f48660c

Observation ee780b3c-0826-449c-ab4c-ea88790560ba · outbound

This paper cites AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.671246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.671246Z digest=sha256:d6250302c80cd794db96aa19053f5bc7f3c50e010ed40866ca31b3feeaa575f3

Observation 8214eded-d1f5-4045-a1ef-9b3c8ccda58b · outbound

This paper cites Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.675859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.675859Z digest=sha256:58861d87ad4f3e639af052a9717338a9ab8359f6077c87089e16a2f7179039f5

Observation 8081f179-da4f-4f01-9f50-a433654c7da6 · outbound

This paper cites , author Xu, N.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Xu, N

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.413897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.681800Z digest=sha256:6a11dd9773fe48ba5d4f2443478ef03e4862ba1dfac7ee582db3804ba4bbb20a

Observation c1595a31-3a4c-489e-9d57-1e06d74e2ae0 · outbound

This paper cites , author Feng, Z.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Feng, Z

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.400845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.686421Z digest=sha256:84e536e4429264524add71044331e701cc0b79435f66bb7f84235ca5fe345e81

Observation 63b0eb18-1d74-4ad7-8018-a6f9b920b10e · outbound

This paper cites , author Phan, L.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Phan, L

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.387649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.691155Z digest=sha256:2a9843c650f81fe8313b9a8c41ad640931ca879757433f19cd9b19a5bb399803

Observation 9e20dac0-cba0-46c5-872f-f3e9580ec8f9 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.696363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.696363Z digest=sha256:caa314ee7d0be309a0cc520d05715b621ab2f9c660eb78781f90a1393899f767

Observation c8102a52-8326-4d43-a705-e939f9ecd07a · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Steering Llama 2 via Contrastive Activation Addition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.701320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.701320Z digest=sha256:19cdb7a21f83b621d18b04af35d8b57d3066992cd213c5f2e2127b7750729de4

Observation 6607a1e4-649e-4bdf-9212-e6e75f0cdb88 · outbound

This paper cites , author Wong, E.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Wong, E

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.374326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.706116Z digest=sha256:7fb3c201fcd50b44fd86ba557790a0dde4411b039793606c6f9d664fb8856fa7

Observation 04604d75-cde7-4c1e-a00b-1aff869832e0 · outbound

This paper cites Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.711503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.711503Z digest=sha256:0558c2858ec04bfb6510a2a9a348d01f9205a8287a5730d7aa218892788630b7

Observation 2b769401-6465-427b-8405-f0c8a8162dec · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.716089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.716089Z digest=sha256:5c81a3132a2ced21b7ed607f990b7ac21aaa71f372576a699702d2bf7e55acf0

Observation f0417e44-1297-491b-ab59-55f0d9698a60 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.720455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.720455Z digest=sha256:b4d82c7653e84121cbc70c8c5c7300c31fbbe07b53200f86fa6adac862fe3ac3

Observation 181e5b86-7bd9-43df-8c52-5359bc7c1512 · outbound

This paper cites , author Chen, K.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Chen, K

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.359315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.724684Z digest=sha256:7afb234d3bd517468133a561dda7a5eac3a141da9c73c158e5321f8d43772f35

Observation ca1e0899-d08a-4dda-983b-41757f54af9f · outbound

This paper cites Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.728950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.728950Z digest=sha256:766c8562707722460aa2bf75ea54effa0b6d108405b8082d0a96eeca4b88662d

Observation 049bbf7f-8aca-4712-b114-6937e92cb423 · outbound

This paper cites , author Haghtalab, N.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Haghtalab, N

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.343484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.733123Z digest=sha256:7f50365a4b9e371f2be14949383ec48f99af2b13f2c2d592314b6487ad70b4f9

Observation 19cf3e0d-9de7-48fe-b988-f87b47517819 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.737120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.737120Z digest=sha256:104e1f201c1bdd0bc731064450b1beba4ca01b713bfd4e4ba8d843bc3b88a6ae

Observation 9163df2c-75da-4783-9748-385de29a21c1 · outbound

This paper cites , author Yi, J.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Yi, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.328209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.741382Z digest=sha256:572d1dbe59fe2b4ad16255a91c7e99f07689cf52fc37b638161461e3c3f84fba

Observation 24072a27-516c-470f-9abc-3e840554bcff · outbound

This paper cites , author Huang, R.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Huang, R

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.312988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.745608Z digest=sha256:fdb695e7956c8d55506b0345216c1b46cc2b7209fe23d8c1905cc77224641dbf

Observation 0c14d0ef-7e00-433e-b8f6-874c769a2d98 · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.749591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.749591Z digest=sha256:b43e1e317a89a5ce204ccef50f73a2bceb6f3e80a52834c61d07fd29bd50fbf4

Observation b290920e-109b-4cc1-83dc-4f29b1d9f7a4 · outbound

This paper cites , author Ye, R.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ye, R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.298063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.754352Z digest=sha256:375d63e002ea7d0100ca9aea31b947086d837aa4e62ea04395de3744d87566ac

Observation b5d3cf62-d8ef-4013-b15c-4e685fa11192 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.758667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.758667Z digest=sha256:6ec2ecc5207e3915dc2d6cd19bb9a22952944b6a427dda6de28f659d45a2f690

Observation 4fee971d-24aa-41bf-8321-ddf3da8abf56 · outbound

This paper cites Robust LLM safeguarding via refusal feature adversarial training.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Robust LLM safeguarding via refusal feature adversarial training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.764576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.764576Z digest=sha256:efd5aa3e3cf4352cb7bdb6d4878f83ec241ed6d4facefd7895af99c206dfcbbf

Observation f76712e0-a05a-4230-92d8-9c76d534f4a9 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.769602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.769602Z digest=sha256:e131efba44176076d4b5f2f42b5203d9c0f119ffacf6cc59e40d8b7f897e011f

Observation 10e3f891-9a1d-4105-ba86-3c0a9a86e5eb · outbound

This paper cites Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.773900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.773900Z digest=sha256:f7bf3ec30fcc581401537db6f09e61bb683b762d31be83afd915b9e0a6842bc4

Observation 4469a48d-a05d-45a0-a860-e2e544ac5b45 · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks On Prompt-Driven Safeguarding for Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.778736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.778736Z digest=sha256:421a381c6b279ab8ac1d98e3fb5ca769efd039a542a0bea18d2dee7029c78fe9

Observation 38181aea-7c25-495f-89aa-4591513365cf · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.783546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.783546Z digest=sha256:41978e3c251512e5495a72b09859460090147021ffa61b78d04f47f081718ff1

Observation 45c68195-138c-429a-ad67-aac2c22e6de5 · outbound

This paper cites , author Phan, L.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Phan, L

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.282597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.791055Z digest=sha256:2283230ea11a04aa1990801025da184738d8bd8ddd3c4648341cc0a3f1266731

Observation 0d93d1b8-eeda-45b8-a0b4-ecfec518d5c7 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.795490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.795490Z digest=sha256:ca249dc15a34e7bf2a74553895e0ff036062aa8ccea2cf3e4ab480951aaf147c

Observation 259f26f1-f9ba-4b42-b94c-8491424e6f70 · outbound

This paper cites write newline.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks write newline

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.799775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.799775Z digest=sha256:9f9c43276431ad6a9457cae29ff7612a948a88f24edd2c83e3654fa53f1ccc09

Pith citing papers

Observation ee1a0b98-dd8a-45a9-b9d0-7bf12bcd306d · inbound

Efficient Safety Alignment of Language Models via Latent Personality Traits cites this paper.

Efficient Safety Alignment of Language Models via Latent Personality Traits Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:27:20.107446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-10T15:26:23.290009Z digest=sha256:6252f0b93c710a0617c0f2487588f5ab85cda08c4347d24338afbeca389713a9