Pith. sign in

Paper Citation Record · LEDGER

Adversarial Training for High-Stakes Reliability

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2205.01663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.01663 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:58.812008Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 91962894-06b8-4dcf-b698-faf34f77066a · inbound

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned cites this paper.

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned Adversarial Training for High-Stakes Reliability

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:38:08.643501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:08.362920Z digest=sha256:92080b2679678bc02557d173cb3c79121a5d33252e67f535b47ab66e2a0a17c7

Observation 8ba6c1c8-f416-430f-a4e4-1450569e1f9a · inbound

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints cites this paper.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Adversarial Training for High-Stakes Reliability

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.812008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.812008Z digest=sha256:62673305ee5a84968fcf7fd7742068e90350ec3e7ade30a1bb407a83a9192563

Observation 350a898d-d298-48a8-bd4c-d7a3a13c8b39 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Adversarial Training for High-Stakes Reliability

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.617525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.617525Z digest=sha256:413a7260ba0b18b46829bf85ab4ceac54b9df77902d9866b4787435f8db324b7

Observation 9e946724-1d6a-434f-9614-a098a67d34ce · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Adversarial Training for High-Stakes Reliability

Reference 278

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:54.706634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:4ab127b8d4dc1bc4ac9eaa96d084a018f97f533cfc7b5fb7eefb1db6eba0ec29

Observation 30b11f2d-32db-41f4-a69a-f0936bd6d3e7 · inbound

Safe Inference-Time Alignment via Lagrangian Reward Augmentation cites this paper.

Safe Inference-Time Alignment via Lagrangian Reward Augmentation Adversarial Training for High-Stakes Reliability

Reference 103

Resolution
unresolved
no resolver link, observed 2026-07-12T07:05:47.150308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:05:47.150308Z digest=sha256:e42708ae17577d1941e60827e76fa4bc071bc7c96f8aa7e4c78aaffad97483a7