Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:35:48.626472Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2507.14202.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:35:48.626472Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
94 of 94 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5b6a9238-00de-4972-a6c5-c256e98d095c · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25cf31f4-6c6d-409f-b5cb-f7b395d031a2 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2af6a0-2e86-4504-8cbc-6ffeed6a1cde · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb252c79-5804-47f0-aa0b-2b0227ac2c84 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Generating Natural Language Adversarial Examples
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffcb86b7-123f-42ce-8130-0269d61955ec · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training A General Language Assistant as a Laboratory for Alignment
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adcb3d4e-4952-4233-a7a6-2fef21cd219c · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f8b8625-ea52-45ea-9324-d027089e80ba · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b29ff979-9448-4e35-8bf3-47d149d4744a · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Constitutional AI: Harmlessness from AI Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517e308b-7e38-4685-85ce-e9f649f6d4c3 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Image Hijacks: Adversarial Images can Control Generative Models at Runtime
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d51cbcab-b7d2-41ed-9e8c-b1dc9bbbc282 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8202b0ae-a407-4800-8fcc-ad004b22a9d0 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21677c54-689c-4f87-803a-53f3d739d648 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3d1067-2d5b-4af6-bd46-b95bde12f133 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training On the Opportunities and Risks of Foundation Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bafa7bb3-a8c0-48c3-884f-7ca91f77b74e · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26876fb-8ead-4a31-9765-0039acb31008 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb4e62a-f278-4460-86c3-e2119a30d3e8 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5a0256-0ed2-4a11-a324-0a6ae5dfe732 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee47c9ff-dcb0-43a9-a097-9fad4e0e75d8 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd144ea-56ba-462c-b157-c2c2fccc610b · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f3b1b5c-6c78-4d68-a448-028cd001b60f · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 670b6015-0d51-4533-b7e0-7978b003fe56 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating Large Language Models Trained on Code
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aa09c3e-12c8-4063-9d99-bb9690b6accf · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training PaLM: Scaling Language Modeling with Pathways
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d70f730-9dc9-4af2-bf75-7a0924bd6251 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11d8663-aa43-4de8-bd1e-a858946f0824 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Supervising strong learners by amplifying weak experts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f0d440a-ef33-4ceb-9f85-df8526eb4c94 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c904d3f8-3630-4b9d-8b71-f52597e821f0 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Verifiers to Solve Math Word Problems
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6639d9f-321d-44b7-ab7f-ad5bad9c5d6b · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Multilingual Jailbreak Challenges in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c61a06-cfae-48e7-abb9-2bb1571c8feb · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9acddfb3-9cda-4059-94d5-cd8bd0434b18 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Understanding parameter differences between analyses employing nested data subsets
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f34939f-ae70-482c-8392-47ddccbdb6f8 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c25ec9-a7aa-48cb-96d2-11d2f2c0cce6 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training HotFlip: White-Box Adversarial Examples for Text Classification
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f696111-d4ca-4c96-85f0-7d428d2af8ab · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42d8818-3e1d-410f-87ba-4fa7f898a85c · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce2fcbd-d0d9-4e9e-9491-0c12d0b58c02 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scaling Laws for Reward Model Overoptimization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90ddecf-99c1-405b-be46-b442ad14d8ed · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebcc7130-ae12-485b-8d12-beb2658a661d · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explaining and Harnessing Adversarial Examples
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b94e92c5-e06a-46e3-8ddb-3c4b945c6be3 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590e6d0d-c3b9-4aa1-91cb-f62b13467cb5 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0358867b-52e6-45ba-b59c-97ede6e52fe3 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Measuring Massive Multitask Language Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28085756-3d06-4400-9799-143a92401037 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Compute-Optimal Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df30fc5-7a42-4e0a-883e-420551bc9683 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training AI safety via debate
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3918a90-c15b-44a3-8567-3f9d1bed0a53 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e4b1b4f-1cda-4008-9ac4-a6828bb2fbb5 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8483a57a-aedc-4071-8651-da0a3aacd10c · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72828130-3626-46ec-aaa7-9b4442706d5a · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Automatically Auditing Large Language Models via Discrete Optimization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5bdc0bf-891b-4c3d-9cf6-b5b29d6825b6 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Single-Source Shortest Paths with Negative Real Weights in $\tilde{O}(mn^{8/9})$ Time
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e39d5a0-df53-4a41-813c-fd6d04b2881a · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6893840f-8913-456c-8bfe-03992331d013 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable agent alignment via reward modeling: a research direction
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2709b8d-4f65-4d21-ab92-6b9473858fa6 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training BERT-ATTACK: Adversarial Attack Against BERT Using BERT
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd64e832-2a45-4da9-bb99-41539ed0ca24 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Let's Verify Step by Step
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b907c4eb-d916-4cbd-8fa8-60a699c361d2 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ea4d27-d6a9-4d67-a875-6c9bd7c3dada · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8382987-8b29-4efe-b8c6-1a8ea5283f86 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Towards Deep Learning Models Resistant to Adversarial Attacks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 455d04b7-778b-4caa-897a-b06d86de4b1b · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53fcdf3a-c28f-4d2c-89ca-8008a48a9534 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77b7dbcc-5efa-42cb-92e5-75aeb0e2ccae · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Teaching language models to support answers with verified quotes
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55775cee-9169-4500-90d2-4eb9a2b99214 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5cccaa2-679e-4f61-81c4-4e13dd8d6e5f · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 103ae790-ec3e-4edf-bf4e-122cf26e8719 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 482fea23-f239-4910-9841-c2975ba97d87 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training WebGPT: Browser-assisted question-answering with human feedback
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a701af0a-b1cc-4a10-8629-972734d4f80b · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable Extraction of Training Data from (Production) Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 569713c6-6f4f-406f-9e2d-46aca8eaa3e5 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206bfffb-3393-4296-b8ed-1a4e1c1631e4 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models with Language Models
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1811695b-06f0-464a-b606-483a2f9e0eb1 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Discovering Language Model Behaviors with Model-Written Evaluations
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8e7847-6a86-4fb4-8f3c-f5a359046bd3 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ignore Previous Prompt: Attack Techniques For Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83499b1b-c548-4c87-9a07-9c519391f9b0 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3fb3cb18-7a56-41ee-aad5-997006c0baf9 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training Can Hurt Generalization
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 453ce8c0-9e8e-4e86-8651-b26a2615ab74 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1664621-df1b-4744-89eb-98e8f0ef31ba · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Proximal Policy Optimization Algorithms
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b853babb-024f-485f-bda8-0d6e9d553ba5 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09e47225-07a6-48db-9c59-e42421f41fe9 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0580d262-b859-4833-9516-6a4889a99700 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4360779b-94b6-4e1e-9494-feb52167a0a3 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5123c869-efcf-4be5-b141-c88b49df9fec · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ensemble Adversarial Training: Attacks and Defenses
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7629bdb-6602-453e-9138-b15b6227f1d3 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Robustness May Be at Odds with Accuracy
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefb1499-3f67-469c-b28d-49a7afe5f42f · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Solving math word problems with process- and outcome-based feedback
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923ae7d5-0634-43e6-9a9a-5af5f2a7d9bd · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal Adversarial Triggers for Attacking and Analyzing NLP
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1581c38e-115f-4014-b61b-b00e5933eef6 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c204d952-f567-4cbe-8f4a-21ee9e541aa8 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6314ad03-312d-443c-ada9-6fc5fa4f6b2f · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Natural Language Adversarial Defense through Synonym Encoding
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db3733e-7e4a-4a0e-ad85-2382dfd528e3 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d9b036-4365-4a9c-b6c8-01475b3ad402 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85d265f5-a6dd-4f4c-bada-8c01f9655ca4 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbroken: How Does LLM Safety Training Fail?
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822f62ee-5f4a-449b-8f91-4f6dc81ff76e · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ethical and social risks of harm from Language Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c54eaadb-dc35-4b7f-a38a-b6c517553455 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Exploring The Landscape of Distributional Robustness for Question Answering Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d55785-7081-4caa-a72b-e802dcad7c97 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72c27dd0-14f0-497b-baa6-8cddeb7a5a7a · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training ReAct: Synergizing Reasoning and Acting in Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 015af25f-f1e0-4a56-bfe7-fea6ed1beb49 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Low-Resource Languages Jailbreak GPT-4
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84165dfd-1663-421f-965b-d33159b6d948 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88de92c1-1199-4988-946c-6ef565782d31 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b547ac2-a353-4344-a0a0-525fd1af06e9 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for Large Neural Language Models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e6d12c9-8c1c-48d0-a294-1eeebdd0ef52 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training FreeLB: Enhanced Adversarial Training for Natural Language Understanding
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 350a898d-d298-48a8-bd4c-d7a3a13c8b39 · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for High-Stakes Reliability
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba77999-6c87-41b9-bf16-cc2d7c3a81ad · outbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.