Pith. sign in

Paper Citation Record · LEDGER

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

As of 23 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2605.21602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21602 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T17:17:39.899234Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T09:29:09.165563Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact26
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6ec32319-cec3-4fae-9162-04a616455d31 · outbound

This paper cites write newline.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs write newline

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.755323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:9f5d2b45305da9d855dc35f2cccaa1d816b93f6c031d2244be20753269cb53f3

Observation 0f438cc6-c8ee-4902-8eec-eef04dfba3ca · outbound

This paper cites Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.068019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:60ca303da92e2eace4c46a9a26512f188deec87c300a9f7e62537e7dca5f5331

Observation a002c8d8-baa1-42a8-a1c3-35d1d79aa95e · outbound

This paper cites System card: Claude opus 4.5.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs System card: Claude opus 4.5

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.753766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:fc205a2a7c9c0766fb82f2e4c9782137d39fb6dde0bcbc7edc01371bb79e1066

Observation c214ed0b-50a6-41e2-a30a-ddaf7fa91f3c · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.070495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:0471af71d13ef38794a7a8c4c6b94fee18e31b422a4c4e5646022e6bd9619520

Observation c002559e-0bac-4742-ab05-ab59cd6c17f1 · outbound

This paper cites Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs , May 2025.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs , May 2025

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.075745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:fb3085e37795096d752656f26f4d2dfb20941b74a273b7b0cab28cd9a407237b

Observation 16e58473-8066-4de9-81e1-91f397ba4aa0 · outbound

This paper cites Envisioning Outlier Exposure by Large Language Models for Out-of-Distribution Detection.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Envisioning Outlier Exposure by Large Language Models for Out-of-Distribution Detection

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.078462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:4131ecdccff352ed23f61f7f615ff703c7770ed4cee924eccffd0c59c002126f

Observation 40d4c4cf-b8b7-47ea-9721-90088215b7c9 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T17:34:58.051793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:f463464b0bf7e0e73f23977ef6b60ae8d19d546cabe52ea14c48402be9f46484

Observation b15210eb-531f-481a-a4fa-d09391c34b4c · outbound

This paper cites LLM Jailbreak Detection for (Almost) Free! , url=.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs LLM Jailbreak Detection for (Almost) Free! , url=

Reference 8

Resolution
verified exact
doi, observed 2026-06-30T17:24:56.700011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:303c31d98ff5fed7422df52db84fecf9899e7e1c734c906ea9a9256655bd2dc9

Observation b032ec18-6c5a-4744-ac3d-e20798e1e329 · outbound

This paper cites Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks

Reference 9

Resolution
verified exact
doi, observed 2026-06-30T17:24:56.691394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:2c6ebf1085d1f5774ce0a02a005c4f50561b1402625551a851002a0cf0cd607a

Observation 9188fab8-dca6-4f40-a9e0-bf45adca1527 · outbound

This paper cites Investigating truthfulness in a pre-release o3 model, April 2025.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Investigating truthfulness in a pre-release o3 model, April 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.757295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:275e0d7fcef6003b42edb8f51ed4d245a06aeb0f1f6b5194ddebb750a6a741cd

Observation c0c514f6-1221-492d-82a9-150c6bb749d0 · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Reward Model Ensembles Help Mitigate Overoptimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.057414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:1544659790e0b488794dbeb7ef76d5aff6a6ffe98b21e47c10caed21bcd234e4

Observation 76842f9b-559e-43ae-8c8d-84afba033d7b · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.060030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:741cbee16da2c2f4654adf7eab278fd6e508815dbebf59c0145311cd2a053ae7

Observation 730f4b53-9c12-486d-bc94-a10b5a409686 · outbound

This paper cites Exploring the Limits of Out -of- Distribution Detection.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Exploring the Limits of Out -of- Distribution Detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.759107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:5cb201c3d1fea66d5eafe3efb566fecf9ad10e1a57a458b9ad1dc2226fd48a18

Observation 4a2875d8-92ac-49a5-bd89-6eb2d8742e5a · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.056745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:12c1272f74c41138c72efbb1197e1a74eff8d01bcfc941a865e018d0ac92472c

Observation d656cc4f-ed68-489c-8ce5-a74ceabaa8be · outbound

This paper cites Alignment faking in large language models.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Alignment faking in large language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.022453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:54daaab5a894c057de9c8572ac4880ca14358c75548a5861674ffd66ad3f604f

Observation 66c88f91-39cc-4eac-ba01-86ab57b81dab · outbound

This paper cites A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.072959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:a7ef15a024e010622bd84714ebf202c393ecc820538e405d249a08ae7f243f0a

Observation 5a5acb9d-442a-471b-b721-537d64ad6e37 · outbound

This paper cites AI Induced Psychosis : A shallow investigation.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs AI Induced Psychosis : A shallow investigation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.752108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:b110e9160218fbe126ecaa2ec56dbe4f1614753b9009ca356ce95902ab2814dd

Observation 3be79a13-9eda-4153-98cc-ecf66d727097 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.052003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:b41d8b6d1ad34a0f118ababd1b4a891f5d2da8fac772c0ae2735a8054fc9077b

Observation 1a782b71-ae39-4aa3-a6ed-6f9cbbfd868a · outbound

This paper cites H idden D etect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs H idden D etect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States

Reference 19

Resolution
verified exact
doi, observed 2026-06-30T17:24:56.695486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:533c73f8ae5fd1c80184ee20031b00b457f9afca1737c0a09c21190e382f7ee5

Observation d65e3e45-9849-4cb5-b7b0-ac93761c47f6 · outbound

This paper cites P., Fishburne, J., Robert P., R., Richard L., C., and Brad S.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs P., Fishburne, J., Robert P., R., Richard L., C., and Brad S

Reference 20

Resolution
verified exact
doi, observed 2026-06-30T17:24:56.693264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:d7e5899a800a6c615f0ec49e543feed99100d3d4735fcebc1d4d6d9100cede02

Observation aa3ffe47-12d4-402c-99f1-17afa0409702 · outbound

This paper cites R., Marks, S., Leike, J., Askell, A., Olah, C., Hubinger, E., and Price, S.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs R., Marks, S., Leike, J., Askell, A., Olah, C., Hubinger, E., and Price, S

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.744602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:eaa69c94418c6b307bb3bd821074dc2fb477027011ee9d0766a859506d180ae6

Observation 863a4e23-0333-471b-810f-f5af3f88e41d · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.054373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:e9b3cf9e3c97c74ae059c8eb100e603591a4fad48b1192a4906547d9f1cf8910

Observation c925d885-78b0-417d-a51f-47b02aa2d92a · outbound

This paper cites A Simple Unified Framework for Detecting Out -of- Distribution Samples and Adversarial Attacks.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs A Simple Unified Framework for Detecting Out -of- Distribution Samples and Adversarial Attacks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.746502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:1609762441d01ccf885f409700c928a3cd97057b8c6813e6c122b0bc76ab6033

Observation 4dae1c17-110a-4568-a574-baf455785c07 · outbound

This paper cites Learning to Detect Unseen Jailbreak Attacks in Large Vision - Language Models , January 2026.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Learning to Detect Unseen Jailbreak Attacks in Large Vision - Language Models , January 2026

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.046934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:4596c2a974d6947865e455fee6743a8bc8398aea58c06b3f860ba3701e16d71f

Observation 1642c00b-5842-448c-a896-0535b5c8d987 · outbound

This paper cites K., Ritchie, S.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs K., Ritchie, S

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.748224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:049d10a47eacc61c2b780c2d49509a657bae7f580077fd61dcaa19f4cefecebb

Observation d0e3c464-3155-4460-b8e8-08b70f7b8cac · outbound

This paper cites an unresolved cited work.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-07-08T09:34:48.741018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:11ceb1402be38bc4af76236e606416af03343c59b57972e05bd4a032bf4c4d22

Observation 0fdb3ce2-b2f6-4fa2-9102-54134080f41a · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Frontier Models are Capable of In-context Scheming

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.041151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:8e813d43d90b665c84ebe5cd358bfbfef8499c45df94be0ef284c0671281b4b9

Observation 704af7ce-51b5-430e-b2fa-5c682f949d79 · outbound

This paper cites JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.046775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:f88ceb68f3b9749c9c25a49ef08b9f3a5541e7f386f4329a01cbbf67b9266685

Observation 903a2002-4550-4d5c-b149-1655b7ac88f7 · outbound

This paper cites Technical report: Performance and baseline evaluations of gpt-oss-safeguard-120b and gpt-oss-safeguard-20b.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Technical report: Performance and baseline evaluations of gpt-oss-safeguard-120b and gpt-oss-safeguard-20b

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.737462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:9a1e756e985126671eb4fc93068c8494f9b5e22601afd156c419554e671f65d0

Observation 296a9e47-0df0-438b-8ba4-33282ff822eb · outbound

This paper cites GPT -5 System Card.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs GPT -5 System Card

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.735639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:f2176c6065425e1cd9f4f7ab527bca0aae718a73cdf048446bb16b903f444079

Observation bbb1b809-2d62-4612-8245-eeb3ed3f0a2d · outbound

This paper cites Sycophancy in GPT -4o: What happened and what we’re doing about it, April 2025 c.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Sycophancy in GPT -4o: What happened and what we’re doing about it, April 2025 c

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.739310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:143c33b793bf8722d5fdb2cb7266d08e2411d50f7415f8e1f7bdf0b64300e526

Observation dac022e5-8dc4-4c02-8757-7573878f4819 · outbound

This paper cites Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.032978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:82a1fa094b9329aedb9d3a58015dc2bf5b6962b7b4b5ef04ed092879dd1b0830

Observation 0758f89d-b92c-4f6c-a22e-85d935a18afe · outbound

This paper cites LLM s know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs LLM s know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts

Reference 33

Resolution
verified exact
doi, observed 2026-06-30T17:24:56.698114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:54315d9bd81f553e194387a0fbdae90b7c4297e985d3d9412edc09a1346186d5

Observation 2e4988e1-8396-41cd-ab28-079532513343 · outbound

This paper cites A Conversation With Bing ’s Chatbot Left Me Deeply Unsettled.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs A Conversation With Bing ’s Chatbot Left Me Deeply Unsettled

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.742843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:a91d54ff21b1724d3ad3f6f3b169492538484458b539b54247dcb04151388c5b

Observation 102de84e-aab2-4395-9730-832ff66ed5b0 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Towards Understanding Sycophancy in Language Models

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T17:34:58.030332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:d4e327697e5f864caf9f558bc12ca01d41d0f16bd6f76dd476072e1a54a4bf64

Observation 4960b4d3-97d6-4421-98f9-e0414b51b9e7 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.062710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:778cbd5665e88eed988dba331958302a2f5c9430cb71cd0cd934512c698b4232

Observation ec2edf4a-df9c-4d08-958f-4d5daec60e24 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs A StrongREJECT for Empty Jailbreaks

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.008198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:42c16344d30f10acd1b90343a72774ded252c7fee0fe447c0be8f7f764b39221

Observation 0808bdf4-e277-479f-8189-f5513f430c8c · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail ? Advances in Neural Information Processing Systems, 36: 0 80079--80110, December 2023.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Jailbroken: How Does LLM Safety Training Fail ? Advances in Neural Information Processing Systems, 36: 0 80079--80110, December 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T09:34:48.750199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:fddbfd518ba3d30abc793bf391033dbe6509f026d2cdcc956dfafe133758b099

Observation 32ce9421-a677-49fb-bb53-5592957b5e3a · outbound

This paper cites On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.065439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:d3424a316a19f5af13daced59f6a9e6c36a8ec1fb98d9f0ed8e1c99fca974855

Observation 3e377bf9-335e-43c3-a57e-8f6028e00f3e · outbound

This paper cites Large Language Models for Anomaly and Out-of-Distribution Detection: A Survey.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Large Language Models for Anomaly and Out-of-Distribution Detection: A Survey

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.049292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:be911499657c3550b5c6e5b10cb8427d189a7b9acd1e12e7d98596aacd464315

Observation be9bb614-fe4a-464f-abd1-131d3d70294a · outbound

This paper cites month = nov, year =.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs month = nov, year =

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:58.054395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:d67ab4986238890218700fdefc6b4d286c258dfb4f925a991b9c149ac8556f38

Observation aa992ae8-e66a-44c2-bae6-df3ed96d797c · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:58.035560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T17:17:39.899234Z digest=sha256:4d5c42981408d6d8875c52a9f4db7b57c2a8804f5add8577cc247dd7c3c5333e

Pith citing papers

Observation 6582823e-714a-40d9-83de-bec5742298a2 · inbound

Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry cites this paper.

Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-01T09:33:38.712047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-01T09:29:09.165563Z digest=sha256:4e913123af5381a63b93c65b0be027cf451ab9df35782c94fc47d8411f318cca