Pith. sign in

Paper Citation Record · LEDGER

Measuring Safety Alignment Effects in Autonomous Security Agents

As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2605.19722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19722 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T04:24:46.357505Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T22:25:22.293704Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact10
  • verified fuzzy41
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 220f1528-4a7f-49d4-820b-3050b76235a8 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Measuring Safety Alignment Effects in Autonomous Security Agents Christiano, Jan Leike, Tom B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.316205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:c9c8a31db883c272dadede8ce65945c4673114e84ed75510e17ff8f9c906b379

Observation 418ad41f-d931-4834-b4f0-24ca7d35dc73 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Measuring Safety Alignment Effects in Autonomous Security Agents Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.955711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:d9ce4d8cc5d51c924bf883043fa1ebde555149a7b33685cd06dad41001396303

Observation bba8e7d0-162c-4b34-8a43-9543a106f6d9 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano.

Measuring Safety Alignment Effects in Autonomous Security Agents Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.319444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:e2443280395bf2ea3c85cd7836a2a269e5e0216e901de6bb0a70847f3d9991d6

Observation 3804f4c3-570d-400c-a2b2-aad6e6729709 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Measuring Safety Alignment Effects in Autonomous Security Agents A General Language Assistant as a Laboratory for Alignment

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.963243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:4a016b893a16b0fdb7a4a2277bc3293ab22ba0861ccd90118059a6478b667b61

Observation f008e3e8-6c37-4891-828e-426bb8eabf12 · outbound

This paper cites an unresolved cited work.

Measuring Safety Alignment Effects in Autonomous Security Agents Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-20T04:33:23.308097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:950635ca9d324b859c0f667c757ace6e2e3af7544d4c9ddb82bb7d803dc59b59

Observation ec530f60-5bc1-457e-83ff-ee2c48e621fb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Measuring Safety Alignment Effects in Autonomous Security Agents Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.976082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:0b88d57c562bc3604b790fe8b7ca8dc2e3ccfc1890e28e860e93d0b1a0c02d79

Observation 9c3e0a6e-49b3-4230-8e44-b6d404b88b38 · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

Measuring Safety Alignment Effects in Autonomous Security Agents TruthfulQA: Measuring how models mimic human falsehoods

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.322490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:09b25092ffa1a7ac0096a5ee2274323840059ab64f87c8ed1e92212b081b5250

Observation 95ec4f96-4b7a-4caf-924c-c79700b78a70 · outbound

This paper cites Aligning AI with shared human values.

Measuring Safety Alignment Effects in Autonomous Security Agents Aligning AI with shared human values

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.312384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:7dd8f5ae15aa44a9724d8c7012d7f5ef75db62ec8fcc125ec8e5d3d334440058

Observation cda6f5ab-4562-4c1c-80b1-6d8beda9763b · outbound

This paper cites Ethical and social risks of harm from Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents Ethical and social risks of harm from Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.949585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3b04ad2b0fdd9dd3ee1dd254388e39e9d6368d3b70f3864f611d875c323802a9

Observation f964082b-0be8-4452-835d-1d262726894a · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models.

Measuring Safety Alignment Effects in Autonomous Security Agents XSTest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.407220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3354eefab2d6bb3ebee0491a46e4fc3ce07350a5142d8ec5a065becf0409bb34

Observation b9ad7247-074b-4fde-9dc7-fe6fd2c6d444 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.

Measuring Safety Alignment Effects in Autonomous Security Agents HarmBench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.412221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:bc0cc494161559bbf5d5ca705fefc8a173869ccad0e926290db15c1370c8ca36

Observation ebc3ab62-8b90-4211-b0dd-f746002d709b · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Measuring Safety Alignment Effects in Autonomous Security Agents Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.303464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:458251c5b8e035a59f8ab61cc5a1013fadb4d0f1faa5a0cf3621d328c8d897d4

Observation 6e90c35a-427c-4e0b-a388-5e9e9fe3088f · outbound

This paper cites A StrongREJECT for empty jailbreaks.

Measuring Safety Alignment Effects in Autonomous Security Agents A StrongREJECT for empty jailbreaks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.402133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:18e59ae91f633527f7692180a85e470bc06ed94aa75b26653b167816aa053415

Observation 17527ca2-a382-41aa-b45d-139de78f1618 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.969962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:d95cc23aa4744887bb230bb8aa90aa5a514e3424f9d44546b24c7cf12539e590

Observation 2849437a-18b7-4d90-9859-d3229a6cc7b0 · outbound

This paper cites Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems 36.

Measuring Safety Alignment Effects in Autonomous Security Agents Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems 36

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.397419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:13448297fa3b4d69f24749c6708d51bf84d4e3bb8b134f5825d2b83bc8b95957

Observation 33321e15-c800-4725-ace6-6b33edf20776 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.936380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:19492b88b0bad4a440bc7ef1955d5474111885e7650d201d49ac3713854c9095

Observation 1fd1169f-7319-49c8-9434-816def9d1576 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.943272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:6636d3232635f3ae6427845d80ec7721bf052317a2d99a0eb57b6cef217ae3ac

Observation 41639c6d-07f6-4455-9fcc-daca9c01376f · outbound

This paper cites an unresolved cited work.

Measuring Safety Alignment Effects in Autonomous Security Agents Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-20T04:28:07.378273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:5d5271cc64811b629ff2f252d591dc728a32c9005e95e2ae5ca23c90d698b5d8

Observation 76a31408-5d3c-4e96-98d8-fd9c9e15328a · outbound

This paper cites NYU CTF bench: A scalable open-source benchmark dataset for evaluating LLMs in offensive security.

Measuring Safety Alignment Effects in Autonomous Security Agents NYU CTF bench: A scalable open-source benchmark dataset for evaluating LLMs in offensive security

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.369642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:91e9204fab449adff115598e2a289a794297130bd3e1b8a6dce01da57217a2ee

Observation 400283d2-2649-4040-8004-be19f07e0161 · outbound

This paper cites Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique, Karthik R.

Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique, Karthik R

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.374065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:a7bdd35bed2462e7da6715467c5c2340d32058bb45ed5240c4dd87bc5ceccee0

Observation f03c3be7-038e-4f54-a246-c534a2308f11 · outbound

This paper cites SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks.

Measuring Safety Alignment Effects in Autonomous Security Agents SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.383513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:635ce26a9d5483e3f6620960ec3b4b9936b454cad31a8539832ef70985d898c3

Observation 27019421-34fd-42fc-86ab-9158bfcfd9bd · outbound

This paper cites AgentBench: Evaluating LLMs as agents.

Measuring Safety Alignment Effects in Autonomous Security Agents AgentBench: Evaluating LLMs as agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.354803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:967378d09bd6b5fc4ef8910e9c7ab1597df45ddc2e378bddc483e800d5213f48

Observation d3b29e2f-9f92-4800-9812-56596e3d0007 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

Measuring Safety Alignment Effects in Autonomous Security Agents ReAct: Synergizing reasoning and acting in language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.359667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:0ad94e49e5e7a3ca5697e5a19f96bb12d836189e3b95df6103e62ee877337459

Observation 4dfd1d19-94b7-474f-8c4a-e53b2fcb1ece · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Measuring Safety Alignment Effects in Autonomous Security Agents Toolformer: Language models can teach themselves to use tools

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.364045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:cbb536ee166aec064e1965cae7ab12269a4be2a553a034e453391bb6db903529

Observation a8118503-132f-49ff-905c-f419a7d1f340 · outbound

This paper cites WebShop: Towards scalable real-world web interaction with grounded language agents.

Measuring Safety Alignment Effects in Autonomous Security Agents WebShop: Towards scalable real-world web interaction with grounded language agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.336725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3d73f9bac395a3bfc501db6dbd7af90c054103b524baf996140fa65f16012e27

Observation bfb5362c-a5df-471a-9d0b-4dc59079293f · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Measuring Safety Alignment Effects in Autonomous Security Agents Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.345780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:149a29ebd8aeaa932ee4b8c7a5763f162186cf6d03ead071f20374a4760b2900

Observation 661f81f5-df7c-42f5-8cd3-2737a4dc60ba · outbound

This paper cites InterCode: Standard- izing and benchmarking interactive coding with execution feedback.

Measuring Safety Alignment Effects in Autonomous Security Agents InterCode: Standard- izing and benchmarking interactive coding with execution feedback

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.311554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:31eb8748cf2e85b26f45998f7b65e7da5a7671e049d032a556e76957f5858d57

Observation b25b258d-9600-43bd-bc1f-17ddeab392cd · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.306636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:bc9e69b121ae06edf1aa6103d864bfedd5bd7a226c598ac04acf1639c10a63f7

Observation 4e592cfa-8a6e-427b-aab5-108aa9f137f0 · outbound

This paper cites Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press.

Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.316366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:a9bcb6f2a2d9752020cc23dcf95088d282be2a63df590ce22523879b1950a8cb

Observation 1379ec2b-efb3-4963-b406-5bd9979eb19a · outbound

This paper cites Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu.

Measuring Safety Alignment Effects in Autonomous Security Agents Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.350528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:fc841036fa6db3fb3dea237ba3c2ee9b22a5fec68aaa972e3a00eacdf0c5c690

Observation 2a5617d7-1b01-49d4-84b9-6b9a2a614570 · outbound

This paper cites AI Agents That Matter.

Measuring Safety Alignment Effects in Autonomous Security Agents AI Agents That Matter

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.909030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:880a09e946caea8754a1be56e1a0182a50a9b2a9ee8fcc66cd330628cc5c84b4

Observation a0cd6025-45f8-44dc-9a09-b2e67efb4328 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Measuring Safety Alignment Effects in Autonomous Security Agents Refusal in language models is mediated by a single direction

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.388192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:e1d854985c8a1bc9496e1b9b1924a1388f4a75015ad3f2c9bfde13d699f74c0d

Observation f1545e90-6371-4de2-a540-e08bc201e08e · outbound

This paper cites Gemma 4: Byte for byte, the most capable open mod- els.

Measuring Safety Alignment Effects in Autonomous Security Agents Gemma 4: Byte for byte, the most capable open mod- els

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.292643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:c563f769aeb46b5550606e6f8cfe9a1c557153bbe62a845f33c5b1f9008278a4

Observation 2de2b40f-a39c-4519-b79a-c3eadddd3639 · outbound

This paper cites google/gemma-4-31b-it.

Measuring Safety Alignment Effects in Autonomous Security Agents google/gemma-4-31b-it

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.296953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:0dc74c60ecf0f43c7cb25f4f45ff382943668be5e8ccf086a0998b65f1d73a54

Observation 3bbb32eb-4182-49a3-8793-914da6920925 · outbound

This paper cites google/gemma-4-26b-a4b-it.

Measuring Safety Alignment Effects in Autonomous Security Agents google/gemma-4-26b-a4b-it

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.301497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:25f1dafbba8d8915e47387f9ce790e0c70506d80aa1d127b45efca74674c5a11

Observation 3a4d70e4-c6d5-4d12-b1fa-3d08fd5ba7a3 · outbound

This paper cites Gemma 4 uncensored.

Measuring Safety Alignment Effects in Autonomous Security Agents Gemma 4 uncensored

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.392340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:b548a8f52bb271eec530e48d55613320f85dec0f6e6fe7eb4a291ba1e12d5623

Observation 85347c1c-327a-4c5e-be3a-73aaf70318da · outbound

This paper cites Trevorjs/gemma-4-31b-it-uncensored.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-31b-it-uncensored

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.312689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:94b4cc15e96d9f49779f2d112cf98b6dca7ceb7b82edffc53d9bd058aac9adf1

Observation b0b0aa9f-6419-46f4-b3c9-dff5906c0f21 · outbound

This paper cites Trevorjs/gemma-4-26b-a4b-it-uncensored.https://huggingface.co/TrevorJS/ gemma-4-26B-A4B-it-uncensored.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-26b-a4b-it-uncensored.https://huggingface.co/TrevorJS/ gemma-4-26B-A4B-it-uncensored

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.283158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:9baa355c94093256f327ff137f27fd34c208205a44f3ba46385f9b35342c985e

Observation f2cc92d1-e814-4abc-b757-27b5bfc4e98a · outbound

This paper cites unsloth/gemma-4-26b-a4b-it-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents unsloth/gemma-4-26b-a4b-it-gguf

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.287235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:8baf80e8ffb015c00a8c783404d75d3ca44187a9768b8c86044f8cbeb2972676

Observation 03a3de1a-533f-4214-99fb-2bbfdb4b912e · outbound

This paper cites Trevorjs/gemma-4-26b-a4b-it-uncensored-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-26b-a4b-it-uncensored-gguf

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.274516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:628c555f8d27297727587a44c4f4491278666c03d06c7a222c2350e324029e37

Observation 32258fff-5d6c-4ef5-adf5-fec5a139efe4 · outbound

This paper cites Qwen/qwen2.5-coder-7b-instruct-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents Qwen/qwen2.5-coder-7b-instruct-gguf

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.278713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:230d1d2fb0d83334a9538fa98e7fce44d1ca21a756938332a197bbf756ef6a86

Observation c9f19cfb-eeee-4a94-b47e-adc614e8c540 · outbound

This paper cites bartowski/qwen2.5-coder-7b-instruct-abliterated-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/qwen2.5-coder-7b-instruct-abliterated-gguf

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.269782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:34359958bda20cc61981d0648df7c343f131a4692ed2b3667d5acf9fa5251ac7

Observation fc0af708-fba0-4e45-bec9-b62df1096b9f · outbound

This paper cites bartowski/meta-llama-3.1-8b-instruct-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/meta-llama-3.1-8b-instruct-gguf

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.260335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:693bbc6daae10514df204b91b25da9a8c746a672d8e25b786a1f4202f97c5edc

Observation ac746392-720e-46d4-8ae3-6314d678a31d · outbound

This paper cites bartowski/meta-llama-3.1-8b-instruct-abliterated-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/meta-llama-3.1-8b-instruct-abliterated-gguf

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.251233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:6b70114c7c32f1420633de0b7d0f7eb1adb907e2f78e6bd5b0846f11c3f43044

Observation 237585d9-e280-4f38-9eaf-40615192bbf1 · outbound

This paper cites Trevorjs/gemma-4-31b-it-uncensored-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-31b-it-uncensored-gguf

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.255755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:373b806521b8f026de9dae299a86db0fb490046b3e633f7d80fbe3dfa965e80d

Observation a4bc177a-d148-48e1-bb31-53b4ea71c81c · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Measuring Safety Alignment Effects in Autonomous Security Agents Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.888463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:d900a146043b238c5277e52d0f6dd5bd9404e56ae731521bdc9f5da97113cf9a

Observation 5536f233-0c18-4819-8a80-64ee0d9fa500 · outbound

This paper cites The GEM benchmark: Natural language generation, its evaluation and metrics.

Measuring Safety Alignment Effects in Autonomous Security Agents The GEM benchmark: Natural language generation, its evaluation and metrics

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.246895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:4e750ccfd05bac55a00b8d9e247851606bee8e55f76c6af12b92afc68486e41b

Observation c4e7561d-cf84-46dd-b731-288d3f0d31d2 · outbound

This paper cites Inspect AI: Framework for large language model evaluations.

Measuring Safety Alignment Effects in Autonomous Security Agents Inspect AI: Framework for large language model evaluations

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.264756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:31aa1b1f562492ed356eeb7e00d7bcf40b7e6bdd73a2c1a541ae74d524efe1e0

Observation 0672b9c5-6fb0-439c-9e09-65e98af93209 · outbound

This paper cites Evaluating Frontier Models for Dangerous Capabilities.

Measuring Safety Alignment Effects in Autonomous Security Agents Evaluating Frontier Models for Dangerous Capabilities

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.916500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:cec94823f6dca6356b7f8633bf3aa1d4f482c6c49523e75a0abd00930902316a

Observation ce90f471-643a-4f7f-a4ff-9105b12525d2 · outbound

This paper cites Model evaluation for extreme risks.

Measuring Safety Alignment Effects in Autonomous Security Agents Model evaluation for extreme risks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.922999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:ab99d344d586af5ebc23b7fe4ae521554d1d5d5cf8f6ec682ac4f433d50015ce

Observation 529e7e09-8a84-4916-b552-9c5359f675af · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Measuring Safety Alignment Effects in Autonomous Security Agents On the Opportunities and Risks of Foundation Models

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.901772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:94aee4117fbf31a9d9c320be344b04bf6c6d6dd30e261bb615b408f2d4808e0d

Observation 8d4e8e4e-0b82-49b3-913e-fe21b06442b2 · outbound

This paper cites Release Strategies and the Social Impacts of Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents Release Strategies and the Social Impacts of Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.929250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:1c294c8b6aef655a4ba6793ffea6207e20f91eeb78a6aa3cd0e729b2a1c67b01

Observation 23f97ab3-c3d4-405e-9b4e-e7eda5c8e8a3 · outbound

This paper cites Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools.

Measuring Safety Alignment Effects in Autonomous Security Agents Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.895469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:629d1accc808835a2969def770161ba2994da03a9a82ce9b7315f48f9f454b77

Observation 98654d1c-29cb-4718-a0a2-97625e5f0f9e · outbound

This paper cites Common weakness enumeration.

Measuring Safety Alignment Effects in Autonomous Security Agents Common weakness enumeration

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.242343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:9581b37180a5621e3a61442a53d191cd91dc3b4e4751c59a4423515cbc137fa3

Observation fcff3578-61f0-4486-98cb-12060f569178 · outbound

This paper cites OWASP Top 10: The ten most critical web application security risks.

Measuring Safety Alignment Effects in Autonomous Security Agents OWASP Top 10: The ten most critical web application security risks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.228976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3542abe532453d1f1d3bad4a10902b295753efa17be0bf68d0bb9c106d975aff

Observation 926c670f-a1fa-40de-b8e9-c4efa0b9b39d · outbound

This paper cites llama.cpp: LLM inference in C/C++.

Measuring Safety Alignment Effects in Autonomous Security Agents llama.cpp: LLM inference in C/C++

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.233038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:76a0a0c24288839d39b23cd2a90b2ca0aed045ff093acb39088567ce30e92d62

Observation f8821e5b-8e86-492d-9bbc-6d133934d26f · outbound

This paper cites The Hugging Face Hub: Machine learning collaboration platform.

Measuring Safety Alignment Effects in Autonomous Security Agents The Hugging Face Hub: Machine learning collaboration platform

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.237863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:e3580afd47bff803fbd8334a6ad4b0ccbaa1ee4d179c51c1af5bf97e3625e7a5

Observation c9e4f9ae-021d-4d91-a3bc-c29a28ddec44 · outbound

This paper cites JSON Schema draft 2020-12.

Measuring Safety Alignment Effects in Autonomous Security Agents JSON Schema draft 2020-12

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-05-20T04:28:07.224764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:86782c791ba5849c62c9c9b7a19988334354ba71049f5c210e28ea6148c5353c

Pith citing papers

Observation f050463d-cd1d-4e14-9d4e-7d3f342d1085 · inbound

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs cites this paper.

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Measuring Safety Alignment Effects in Autonomous Security Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:22.293704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:25:22.293704Z digest=sha256:29a44bb98c9ffbcb6e7bcf16f5f146ee4cfc471be8cf2ae415cf277f78ebe4a2