Pith. sign in

Paper Citation Record · LEDGER

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

As of 7 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2507.14202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14202 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:35:48.626472Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b6a9238-00de-4972-a6c5-c256e98d095c · outbound

This paper cites online" 'onlinestring :=.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.238850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.238850Z digest=sha256:13a23acfe71cfefbc512d30e31a12b807b385fdf67420af6b61efe175e188721

Observation 25cf31f4-6c6d-409f-b5cb-f7b395d031a2 · outbound

This paper cites write newline.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.390160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.390160Z digest=sha256:b1dcd2d2202ce0f3130f3a38e7980705ff43f4080cd06f764393993d1a864935

Observation 5a2af6a0-2e86-4504-8cbc-6ffeed6a1cde · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.424574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.424574Z digest=sha256:d314f5e474151e328f45e0b44084b0d1d47d3d0c231fc0a6cc6bd846d69344f8

Observation bb252c79-5804-47f0-aa0b-2b0227ac2c84 · outbound

This paper cites Generating Natural Language Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Generating Natural Language Adversarial Examples

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.437181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.437181Z digest=sha256:477c1dbc314470eb7371dcac4b623375b61456d59a5ab2a827f52b07f5f34538

Observation ffcb86b7-123f-42ce-8130-0269d61955ec · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training A General Language Assistant as a Laboratory for Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.468119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.468119Z digest=sha256:c667c461652c5720261cd396ae361d3860f09bd1d4861bf9207896cf1ecaa809

Observation adcb3d4e-4952-4233-a7a6-2fef21cd219c · outbound

This paper cites Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:35:50.534543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:47.522560Z digest=sha256:a48d07aed64501d8b30296348225f1de2267cfefae3ad68eda1152412b6da308

Observation 1f8b8625-ea52-45ea-9324-d027089e80ba · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.608128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.608128Z digest=sha256:f55eba0f4814eb1364214d6da8aebc88e47831e2f5fd88c2f2e6dd31d7cf23aa

Observation b29ff979-9448-4e35-8bf3-47d149d4744a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Constitutional AI: Harmlessness from AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.723342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.723342Z digest=sha256:604d4d7edd948c329235b718a2569b6bc0e78ffe6cb66a4173dcaff21dffa14d

Observation 517e308b-7e38-4685-85ce-e9f649f6d4c3 · outbound

This paper cites Image Hijacks: Adversarial Images can Control Generative Models at Runtime.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Image Hijacks: Adversarial Images can Control Generative Models at Runtime

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.822419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.822419Z digest=sha256:9b21b5dec84aedff6d963aca4e5ee15f95bfdbfccb0a7efebd63dd717450703a

Observation d51cbcab-b7d2-41ed-9e8c-b1dc9bbbc282 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.897090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.897090Z digest=sha256:67219d0e2b0d3f017bc80cc8377b9a61826742680caec5a8458fc57b3725d0f4

Observation 8202b0ae-a407-4800-8fcc-ad004b22a9d0 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.960711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.960711Z digest=sha256:6ac1bf508c148442093cc35bda6b4d78537e782a26542795a560f392ee4f589f

Observation 21677c54-689c-4f87-803a-53f3d739d648 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.970414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.970414Z digest=sha256:4fcf9ad80cf8a0439c2010bbb208ec5d3ac112ec7e22ac59d4e51601423a2da0

Observation ce3d1067-2d5b-4af6-bd46-b95bde12f133 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training On the Opportunities and Risks of Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.975977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.975977Z digest=sha256:8247c761559df56cb621c002c6bfb201eb89aa057e0f3a8f218765d0f3d09ac4

Observation bafa7bb3-a8c0-48c3-884f-7ca91f77b74e · outbound

This paper cites Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.982288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.982288Z digest=sha256:cf3a9a64e8e36ca08fa7ce0bca84afb26ae24faa7f29895a684d974c20332a32

Observation b26876fb-8ead-4a31-9765-0039acb31008 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.989197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.989197Z digest=sha256:070b5337857c7d9ebfbf2f8c04def6240af17d3cbb32b10f505ae3752949b0f2

Observation 4bb4e62a-f278-4460-86c3-e2119a30d3e8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:47.996109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:47.996109Z digest=sha256:64fc1b2fa4114c1d4a3bf366ca947d039073ec5c29a1c63b65d16f9072239712

Observation 8a5a0256-0ed2-4a11-a324-0a6ae5dfe732 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.002781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.002781Z digest=sha256:057f0dbc9cdfc24096728f3abdc0da80eb7b02bf588f8f245b90f7f5d8351a59

Observation ee47c9ff-dcb0-43a9-a097-9fad4e0e75d8 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.010807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.010807Z digest=sha256:97a5528807bbcc56ef3f0a3209ece747598b539efca2951460f2edf356e78dd7

Observation 4fd144ea-56ba-462c-b157-c2c2fccc610b · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.018055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.018055Z digest=sha256:1cb17a6fe6e52cd5e19c422a9ba10cc1a7f2e0a2b3a456478f077cd1d53f460b

Observation 2f3b1b5c-6c78-4d68-a448-028cd001b60f · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.025371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.025371Z digest=sha256:322c6b633638bc11c17574a739d7207d76b5c7678f93fcf2113313fb1e124e07

Observation 670b6015-0d51-4533-b7e0-7978b003fe56 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Evaluating Large Language Models Trained on Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.035926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.035926Z digest=sha256:80471b785b1f933a280b949a45b0c4f8b0b47385a0bc76f360df2aa89f226d20

Observation 3aa09c3e-12c8-4063-9d99-bb9690b6accf · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training PaLM: Scaling Language Modeling with Pathways

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.042843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.042843Z digest=sha256:ca425cea0c097810b9fc1cda074431e2cf3cdc69493e2e66deeefb00a456c417

Observation 3d70f730-9dc9-4af2-bf75-7a0924bd6251 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.053465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.053465Z digest=sha256:a97bc9c300b1aa39808e04eb6f0d9c52f657a3aeedecf56a7dd2c94eba51a06c

Observation a11d8663-aa43-4de8-bd1e-a858946f0824 · outbound

This paper cites Supervising strong learners by amplifying weak experts.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Supervising strong learners by amplifying weak experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.063524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.063524Z digest=sha256:26de992d40112da7be6ec02543dfe0c348a6dab4406c2d7c01f69fa5e77068c7

Observation 9f0d440a-ef33-4ceb-9f85-df8526eb4c94 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.070951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.070951Z digest=sha256:132d715e84b56a80ba7ed7b4d3c24f890eaf27588483c92cdc590bbedb7aff41

Observation c904d3f8-3630-4b9d-8b71-f52597e821f0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Verifiers to Solve Math Word Problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.078087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.078087Z digest=sha256:6a3dd538b879011d5f9850856aad3b3015b60d7c11d2319b6243543b6778861c

Observation d6639d9f-321d-44b7-ab7f-ad5bad9c5d6b · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Multilingual Jailbreak Challenges in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.084454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.084454Z digest=sha256:7bf642376601de627c38b31cb6c632dc92523ec16899ee34e05c02914a5c28e0

Observation 68c61a06-cfae-48e7-abb9-2bb1571c8feb · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.092363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.092363Z digest=sha256:7fcab70e35b0ca477279ad7abb5895f3c9b76e97ebfe3c67a3392630b5881cc4

Observation 9acddfb3-9cda-4059-94d5-cd8bd0434b18 · outbound

This paper cites Understanding parameter differences between analyses employing nested data subsets.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Understanding parameter differences between analyses employing nested data subsets

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.100821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.100821Z digest=sha256:edb0a23f2e948d17ea9c5e688f1371286f9d6e30c13e31ef9b8995a8793d6233

Observation 7f34939f-ae70-482c-8392-47ddccbdb6f8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.109664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.109664Z digest=sha256:9b05116b9ef926b97d557eb0ed707ed95a6aad0e152fc689244eaf8365300038

Observation 26c25ec9-a7aa-48cb-96d2-11d2f2c0cce6 · outbound

This paper cites HotFlip: White-Box Adversarial Examples for Text Classification.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training HotFlip: White-Box Adversarial Examples for Text Classification

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.119501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.119501Z digest=sha256:aa74ed566a44ab4a6b8f205048e2df32ed0a8a06c339f18f9729ca0a3f0896c7

Observation 3f696111-d4ca-4c96-85f0-7d428d2af8ab · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.128310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.128310Z digest=sha256:f17392167ec50b0974d7f9549905120cd4407961f6ad6542ff4b58fb5696668a

Observation c42d8818-3e1d-410f-87ba-4fa7f898a85c · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.134541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.134541Z digest=sha256:b5b7aecbc6e9aeb6ff40fde475a7c9f74b3eb049e52204703b6b95d8af6c2ac9

Observation 2ce2fcbd-d0d9-4e9e-9491-0c12d0b58c02 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scaling Laws for Reward Model Overoptimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.142756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.142756Z digest=sha256:4c18c9dcdf1315601c5b11a657241aec944242e008a3cbd9baec3ef6634b175a

Observation a90ddecf-99c1-405b-be46-b442ad14d8ed · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.151517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.151517Z digest=sha256:64a760ee79874ecdb1c4ff18fcd5c57c0bd6c7b0f7209c591d2abd1c1c3dae9e

Observation ebcc7130-ae12-485b-8d12-beb2658a661d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Explaining and Harnessing Adversarial Examples

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.158428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.158428Z digest=sha256:492678ec70aeec584d7b709cb75051189b957bfb594139b79fd02a3152ee0773

Observation b94e92c5-e06a-46e3-8ddb-3c4b945c6be3 · outbound

This paper cites Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.166666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.166666Z digest=sha256:a8d6f59e25658c5e02d53d5cb6aa9c7f63dccb33975d1d4b0921a72a43d4321c

Observation 590e6d0d-c3b9-4aa1-91cb-f62b13467cb5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.981468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.174906Z digest=sha256:fd4d1d9a88f40a78a105d54186b2411324b0d940ddb708496f72d3603f210c94

Observation 0358867b-52e6-45ba-b59c-97ede6e52fe3 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Measuring Massive Multitask Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.184013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.184013Z digest=sha256:ba413116df5b2f6ff06b4693376f9caa38372d024aceee89a6be0ac1f7f0126b

Observation 28085756-3d06-4400-9799-143a92401037 · outbound

This paper cites Training Compute-Optimal Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Training Compute-Optimal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.190750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.190750Z digest=sha256:b8784e16773893f01dd6b7f091d1b0a1b36a78aa471af94e0e88dce1e00a10ed

Observation 6df30fc5-7a42-4e0a-883e-420551bc9683 · outbound

This paper cites AI safety via debate.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training AI safety via debate

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.198213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.198213Z digest=sha256:c93ad92aab8be6c7d4e95a1860ac466d5c41b44fd8c2d9eb52fd24e5363615df

Observation c3918a90-c15b-44a3-8567-3f9d1bed0a53 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.954180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.205695Z digest=sha256:b6853ffead565633f8f1f260de64ddb0107ff5a999d52d97edcab4755de10d6b

Observation 8e4b1b4f-1cda-4008-9ac4-a6828bb2fbb5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.925417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.215145Z digest=sha256:45fe224316ff6d62bcf8ba0201f0eac9234cc30e6b402e07975db29dddd88747

Observation 8483a57a-aedc-4071-8651-da0a3aacd10c · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.898671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.223231Z digest=sha256:f94b4e97853da886587322bca79f94942f684b360dc263e44680838b010132eb

Observation 72828130-3626-46ec-aaa7-9b4442706d5a · outbound

This paper cites Automatically Auditing Large Language Models via Discrete Optimization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Automatically Auditing Large Language Models via Discrete Optimization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.231633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.231633Z digest=sha256:297e82244631cc39d37016b5782d04597896ec09f10732c2df5804e39d68345f

Observation c5bdc0bf-891b-4c3d-9cf6-b5b29d6825b6 · outbound

This paper cites Single-Source Shortest Paths with Negative Real Weights in $\tilde{O}(mn^{8/9})$ Time.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Single-Source Shortest Paths with Negative Real Weights in $\tilde{O}(mn^{8/9})$ Time

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:35:49.762344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.242664Z digest=sha256:24b2d49483285eb11fe47f49ddeebfeab7f36c996ef18ea9373c24785e0950dd

Observation 8e39d5a0-df53-4a41-813c-fd6d04b2881a · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.251121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.251121Z digest=sha256:186ab3a0c5a7169d54808df2e46e2506d1b7da5167e795ed10c5af4794305361

Observation 6893840f-8913-456c-8bfe-03992331d013 · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable agent alignment via reward modeling: a research direction

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.260173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.260173Z digest=sha256:d023a8ffbd548c6c9b75ea8d73112a85baaef97ed1f8ebd60e90947e308044ad

Observation c2709b8d-4f65-4d21-ab92-6b9473858fa6 · outbound

This paper cites BERT-ATTACK: Adversarial Attack Against BERT Using BERT.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training BERT-ATTACK: Adversarial Attack Against BERT Using BERT

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.267389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.267389Z digest=sha256:8d046e2ef2af113642a16eecfa79c0e2c9fff9c668780e8b82bcc14e719871e4

Observation dd64e832-2a45-4da9-bb99-41539ed0ca24 · outbound

This paper cites Let's Verify Step by Step.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Let's Verify Step by Step

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.273778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.273778Z digest=sha256:1ee84c5b9e75fabc779ed9d842b3c20761ff43d24afa5fed3d8a64dc892051ff

Observation b907c4eb-d916-4cbd-8fa8-60a699c361d2 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.280960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.280960Z digest=sha256:1b4ca0aeafdf1338db4a5e40b53d05cb7383bed022257be317572e30551a4c59

Observation 67ea4d27-d6a9-4d67-a875-6c9bd7c3dada · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.289162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.289162Z digest=sha256:c5887da54e7e271ce0acd508ac8761db37a977770c1822864f327f9b6144ad30

Observation c8382987-8b29-4efe-b8c6-1a8ea5283f86 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.302663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.302663Z digest=sha256:6a973f0f5f48708c22bfcac451d91ee5e10b4cb9e9e31cf85c554838ddd5952b

Observation 455d04b7-778b-4caa-897a-b06d86de4b1b · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.871578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.312375Z digest=sha256:717faf8e9d75eb8885fbc393188963f486d0b38e9e961457d55576a7f2d0a0d6

Observation 53fcdf3a-c28f-4d2c-89ca-8008a48a9534 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.318958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.318958Z digest=sha256:e89543746ffc65a2224c71ad9d269b4c9c5f347da6b8807be1642da8018663c2

Observation 77b7dbcc-5efa-42cb-92e5-75aeb0e2ccae · outbound

This paper cites Teaching language models to support answers with verified quotes.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Teaching language models to support answers with verified quotes

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.325823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.325823Z digest=sha256:d97b1781b6215d88a7bfb2c92dfcda262c77e8787d7d049313ce66e82df85989

Observation 55775cee-9169-4500-90d2-4eb9a2b99214 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.842475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.334358Z digest=sha256:6f6956106d8ddd175ca690d208ddd8114fc245c23fb77a49710d158ca81339c6

Observation f5cccaa2-679e-4f61-81c4-4e13dd8d6e5f · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.813928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.340570Z digest=sha256:9cf5a59759e741ad5fb45d61f410980f586a58d3e85b28f57c34c79ab96a695c

Observation 103ae790-ec3e-4edf-bf4e-122cf26e8719 · outbound

This paper cites TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.347629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.347629Z digest=sha256:0af40f928e81ac606584cc8b3565cd0147f0e1c8e0d9eac5e6881006c0f25787

Observation 482fea23-f239-4910-9841-c2975ba97d87 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training WebGPT: Browser-assisted question-answering with human feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.355471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.355471Z digest=sha256:c4341debff8b9640f683a8d311dcaabb4ff5821a5872e14cf96f9db4636c4b0b

Observation a701af0a-b1cc-4a10-8629-972734d4f80b · outbound

This paper cites Scalable Extraction of Training Data from (Production) Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable Extraction of Training Data from (Production) Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.362446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.362446Z digest=sha256:9344866851b657a5e707fd67e27e75ef32aa2e4181ab539ab58760a19e15888e

Observation 569713c6-6f4f-406f-9e2d-46aca8eaa3e5 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.370906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.370906Z digest=sha256:34af6cf9e1ffcc327ba2fbb5ee26b68eb15c9fe5cb04fd60ee0998e921ff73c2

Observation 206bfffb-3393-4296-b8ed-1a4e1c1631e4 · outbound

This paper cites Red Teaming Language Models with Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Red Teaming Language Models with Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.379195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.379195Z digest=sha256:6779ec93fb38715c5bd9fecfce0151731f2dc2e0144f1ae76d4a330391edb03e

Observation 1811695b-06f0-464a-b606-483a2f9e0eb1 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Discovering Language Model Behaviors with Model-Written Evaluations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.387140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.387140Z digest=sha256:ce557a145bff12b761a647e9603c62092faa6447711195e01ae1882ae19038b0

Observation 3b8e7847-6a86-4fb4-8f3c-f5a359046bd3 · outbound

This paper cites Ignore Previous Prompt: Attack Techniques For Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ignore Previous Prompt: Attack Techniques For Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.395357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.395357Z digest=sha256:8c05863c5b225d28eaa69f2324678651d95992dc220b6c249e6a782275808aaf

Observation 83499b1b-c548-4c87-9a07-9c519391f9b0 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.775043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.406503Z digest=sha256:4cfe7ecf90ae125c17096e163d1c2400b980f876f00d2d58176e8ab81d5b2acd

Observation 3fb3cb18-7a56-41ee-aad5-997006c0baf9 · outbound

This paper cites Adversarial Training Can Hurt Generalization.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training Can Hurt Generalization

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.413261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.413261Z digest=sha256:88f0e0a23aaf963dd5b1c4db56da875371f361ed5dee16bbde6fc20d9adc65a7

Observation 453ce8c0-9e8e-4e86-8651-b26a2615ab74 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.752692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.420411Z digest=sha256:bcd8bf0136d933efe726c66d77726118935e151c7ae5b7177a7b6e9aaf70c043

Observation f1664621-df1b-4744-89eb-98e8f0ef31ba · outbound

This paper cites Proximal Policy Optimization Algorithms.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Proximal Policy Optimization Algorithms

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.427520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.427520Z digest=sha256:bb59358a70055f71ad50b65f0642ce3d0aceab3b6d8d5057da571355340face4

Observation b853babb-024f-485f-bda8-0d6e9d553ba5 · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.436360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.436360Z digest=sha256:c085e2cc3c68cf38705924b8918940cd876b3863699e29964ceb4f37ba107577

Observation 09e47225-07a6-48db-9c59-e42421f41fe9 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.445431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.445431Z digest=sha256:3905e81660c470519889e91c3b66311ba3bbcefef6bde77b9d45669702196c0a

Observation 0580d262-b859-4833-9516-6a4889a99700 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.454219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.454219Z digest=sha256:2a044595ea4103121866d8081ffe5a4d33758640b0e2bde717f1517b960508a4

Observation 4360779b-94b6-4e1e-9494-feb52167a0a3 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.710417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.461142Z digest=sha256:9c42014b24cb3acfbd99bfd3377edbd45ea6b22cc711789a1040339ee94c73b1

Observation 5123c869-efcf-4be5-b141-c88b49df9fec · outbound

This paper cites Ensemble Adversarial Training: Attacks and Defenses.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ensemble Adversarial Training: Attacks and Defenses

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.467325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.467325Z digest=sha256:1e3477d31f675a9cf4addc830f5b275d22ab459bfe83248e762b848831743bab

Observation f7629bdb-6602-453e-9138-b15b6227f1d3 · outbound

This paper cites Robustness May Be at Odds with Accuracy.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Robustness May Be at Odds with Accuracy

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.474839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.474839Z digest=sha256:d8361aba3e15c01b76bc3f274af1aa380c6e641f41747076ec11cdbcb0f86bce

Observation aefb1499-3f67-469c-b28d-49a7afe5f42f · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Solving math word problems with process- and outcome-based feedback

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.481517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.481517Z digest=sha256:977981ab3cb1e576090c246ae1ce5476c853d6b060cf21bb2c6cd68f0bc8bb4d

Observation 923ae7d5-0634-43e6-9a9a-5af5f2a7d9bd · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.488605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.488605Z digest=sha256:17bfbbde28e040d921a860e7675ccabcffb2f43b50d714079777cdbf835279f6

Observation 1581c38e-115f-4014-b61b-b00e5933eef6 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.496124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.496124Z digest=sha256:2af6a19afbdffe46b6e8c3bd5fc982c92af9123f7f01c9b60c2fb86e449dd716

Observation c204d952-f567-4cbe-8f4a-21ee9e541aa8 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.688433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.505595Z digest=sha256:4eef32e0ff419d23969a7793339c732f8c9c1f602342689b0433761200d8b1f2

Observation 6314ad03-312d-443c-ada9-6fc5fa4f6b2f · outbound

This paper cites Natural Language Adversarial Defense through Synonym Encoding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Natural Language Adversarial Defense through Synonym Encoding

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.513887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.513887Z digest=sha256:c29a74ed1abd0d51e4d8f7540b40f0aa723ac1a87ef72bc7a5c33fc1ab4f71a5

Observation 0db3733e-7e4a-4a0e-ad85-2382dfd528e3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.522431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.522431Z digest=sha256:0c2fcc2ef5a80e33343acb07c8a11beddbc5bc87245020c6833adb446145da9c

Observation 51d9b036-4365-4a9c-b6c8-01475b3ad402 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.666517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.537349Z digest=sha256:51813b08eef2f05ffa984415959edad1508ef52ccbfdee899e29972c81476af5

Observation 85d265f5-a6dd-4f4c-bada-8c01f9655ca4 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Jailbroken: How Does LLM Safety Training Fail?

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.545520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.545520Z digest=sha256:8c3857aaf0e95ba4ffc3e9138972bb9fd67080118b77819f80ce76393acb600c

Observation 822f62ee-5f4a-449b-8f91-4f6dc81ff76e · outbound

This paper cites Ethical and social risks of harm from Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Ethical and social risks of harm from Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.553979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.553979Z digest=sha256:67c5231d9cfc9a2b9436a60caed1e8501c1f62b98ee007019108d379e8a2dd13

Observation c54eaadb-dc35-4b7f-a38a-b6c517553455 · outbound

This paper cites Exploring The Landscape of Distributional Robustness for Question Answering Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Exploring The Landscape of Distributional Robustness for Question Answering Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.560955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.560955Z digest=sha256:2f4c60dac00c5aeb83c32f2d103d0b2bb79091ce520cb998fd0ca79056b65cc6

Observation a8d55785-7081-4caa-a72b-e802dcad7c97 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.642464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.569962Z digest=sha256:2961e09f52a36e9c8e133e9801913da6147e255199fdcce2fa21a767716de863

Observation 72c27dd0-14f0-497b-baa6-8cddeb7a5a7a · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training ReAct: Synergizing Reasoning and Acting in Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.576790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.576790Z digest=sha256:d6cc8d9f68ca877b2afe17fbd016b89c9add8a88cc06dab74ae006d432b085dc

Observation 015af25f-f1e0-4a56-bfe7-fea6ed1beb49 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Low-Resource Languages Jailbreak GPT-4

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.582594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.582594Z digest=sha256:3ee1a8b3651a0e3750c08bf2f174159b55db4786fba5bc34c7776f116f1e504d

Observation 84165dfd-1663-421f-965b-d33159b6d948 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.591046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.591046Z digest=sha256:1d197ea19ecdf896dc463b923831fd783759331929bf77361c51b502ac2dcbb2

Observation 88de92c1-1199-4988-946c-6ef565782d31 · outbound

This paper cites an unresolved cited work.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:35:50.613453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:35:48.596978Z digest=sha256:9ceaf3b50fa6a6a218080b58e9aee75511ca50f8b2f5651516caa590090b3ebc

Observation 8b547ac2-a353-4344-a0a0-525fd1af06e9 · outbound

This paper cites Adversarial Training for Large Neural Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for Large Neural Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.602842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.602842Z digest=sha256:2f5102be46c1f72a9f113a6aa5e6f1eccbe880b2bb5821be01cd3a8d80ecfc67

Observation 2e6d12c9-8c1c-48d0-a294-1eeebdd0ef52 · outbound

This paper cites FreeLB: Enhanced Adversarial Training for Natural Language Understanding.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training FreeLB: Enhanced Adversarial Training for Natural Language Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.609506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.609506Z digest=sha256:f775008ca84c345cba19fbbcb17cc9e4d16a3235c634b1af85777bee45afa285

Observation 350a898d-d298-48a8-bd4c-d7a3a13c8b39 · outbound

This paper cites Adversarial Training for High-Stakes Reliability.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for High-Stakes Reliability

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.617525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.617525Z digest=sha256:7d8f68996821810c12ae82da15c23e30dd43941031328a05a7e9c314702751f2

Observation bba77999-6c87-41b9-bf16-cc2d7c3a81ad · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.626472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.626472Z digest=sha256:a230168ef072816241fd5279a6bdb345f4bf778a70e09f968b39d81b2efaecd3

Pith citing papers

No inbound Pith citation observations are available.