Pith. sign in

Paper Citation Record · LEDGER

Measuring Safety Alignment Effects in Autonomous Security Agents

As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2605.19722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19722 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T04:24:46.357505Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T22:25:22.293704Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact10
  • verified fuzzy41
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 220f1528-4a7f-49d4-820b-3050b76235a8 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Measuring Safety Alignment Effects in Autonomous Security Agents Christiano, Jan Leike, Tom B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.316205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:16747021a3117af16ad1524839c32bedbb6e1c7def7d93be7a83fe9bd77e9d3d

Observation 418ad41f-d931-4834-b4f0-24ca7d35dc73 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Measuring Safety Alignment Effects in Autonomous Security Agents Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.955711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:831da8c90b9fed37eb46eddbb7081b3130a94237b840b30f47264483e6cf3434

Observation bba8e7d0-162c-4b34-8a43-9543a106f6d9 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano.

Measuring Safety Alignment Effects in Autonomous Security Agents Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.319444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3ae24ace16fa31fe86a28734da0da7933147e7145b19a0b7f0b73f8a932487bd

Observation 3804f4c3-570d-400c-a2b2-aad6e6729709 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Measuring Safety Alignment Effects in Autonomous Security Agents A General Language Assistant as a Laboratory for Alignment

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.963243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:4c41efaca6ac683e9a9a38a579bbe48c4f2f58740a4454c504498a8d8e90753e

Observation f008e3e8-6c37-4891-828e-426bb8eabf12 · outbound

This paper cites an unresolved cited work.

Measuring Safety Alignment Effects in Autonomous Security Agents Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-20T04:33:23.308097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:38e81a839188a645c529784a1e7eeb9099b51c45c11f368f95d1c0a2792219fa

Observation ec530f60-5bc1-457e-83ff-ee2c48e621fb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Measuring Safety Alignment Effects in Autonomous Security Agents Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.976082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:d0c38f4b670674db89f2649fd3a793a695554e3100605552a4d472a59da8a4c4

Observation 9c3e0a6e-49b3-4230-8e44-b6d404b88b38 · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

Measuring Safety Alignment Effects in Autonomous Security Agents TruthfulQA: Measuring how models mimic human falsehoods

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.322490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:37d7507fb926f7cc77cdbe2f2b64bed577bc5e1f4a931055aafd66a0d797959f

Observation 95ec4f96-4b7a-4caf-924c-c79700b78a70 · outbound

This paper cites Aligning AI with shared human values.

Measuring Safety Alignment Effects in Autonomous Security Agents Aligning AI with shared human values

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.312384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:81ce865d12276c57785df3901ed3a4f131194f78686624c365bb33d92a6daf57

Observation cda6f5ab-4562-4c1c-80b1-6d8beda9763b · outbound

This paper cites Ethical and social risks of harm from Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents Ethical and social risks of harm from Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.949585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:8de161d840a7e93bad2805288dd3e75a35e02574c350d20bed4eae3010148402

Observation f964082b-0be8-4452-835d-1d262726894a · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models.

Measuring Safety Alignment Effects in Autonomous Security Agents XSTest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.407220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:b74963975878fc26cc8daa5e1fe32eecab98543128cf494eb189e978a33d81c7

Observation b9ad7247-074b-4fde-9dc7-fe6fd2c6d444 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.

Measuring Safety Alignment Effects in Autonomous Security Agents HarmBench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.412221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:813db8b701f2ea2643aaf6ee36cac95cdc79188e10d39c9450b99ee6200ee24b

Observation ebc3ab62-8b90-4211-b0dd-f746002d709b · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Measuring Safety Alignment Effects in Autonomous Security Agents Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.303464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:00c69234d41a5ba4ae5c919625ddc74ff9b75fa22bd82d0ace54e03fa98a92f7

Observation 6e90c35a-427c-4e0b-a388-5e9e9fe3088f · outbound

This paper cites A StrongREJECT for empty jailbreaks.

Measuring Safety Alignment Effects in Autonomous Security Agents A StrongREJECT for empty jailbreaks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.402133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:6de76070e4ec7f5d24a7faef53efb529770df3ccc97d19223b5113855445c1f7

Observation 17527ca2-a382-41aa-b45d-139de78f1618 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.969962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:9446f237dcdf98bac0e273c0a41d85b1adf6ddf4f114a58ab2c11cb625f06200

Observation 2849437a-18b7-4d90-9859-d3229a6cc7b0 · outbound

This paper cites Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems 36.

Measuring Safety Alignment Effects in Autonomous Security Agents Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems 36

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.397419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:41b34a49bd4bf587ab30a2b15da5aecc01efb54dc81f7b24a8905eae9466130f

Observation 33321e15-c800-4725-ace6-6b33edf20776 · outbound

This paper cites CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.936380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:af59160d4c9428e3fd55daaaacaed0263e1a738b8cb9942eb74b2d632d95230a

Observation 1fd1169f-7319-49c8-9434-816def9d1576 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.943272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:b70a86d8ddc32808678c6d7667164e0e0c237d81e17fc57f530092b0ecdf9bf3

Observation 41639c6d-07f6-4455-9fcc-daca9c01376f · outbound

This paper cites an unresolved cited work.

Measuring Safety Alignment Effects in Autonomous Security Agents Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-20T04:28:07.378273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:60d1aa11b77ce1414f52ebabb7ebd24d7d0257c05e20c503dc9842149c275fa6

Observation 76a31408-5d3c-4e96-98d8-fd9c9e15328a · outbound

This paper cites NYU CTF bench: A scalable open-source benchmark dataset for evaluating LLMs in offensive security.

Measuring Safety Alignment Effects in Autonomous Security Agents NYU CTF bench: A scalable open-source benchmark dataset for evaluating LLMs in offensive security

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.369642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:fd3cb28aa282e377b2e7cf90a76fb8d7df8da48a1b5cd2569179f246be72fd9c

Observation 400283d2-2649-4040-8004-be19f07e0161 · outbound

This paper cites Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique, Karthik R.

Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique, Karthik R

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.374065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:5d83c12e819ce5461241723f82736e4495f36ba2b80e23b33a8a36e909d120b8

Observation f03c3be7-038e-4f54-a246-c534a2308f11 · outbound

This paper cites SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks.

Measuring Safety Alignment Effects in Autonomous Security Agents SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.383513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:d9c6d8ca681f8e4cc6825dcd6b5f6faa34c1c52faf3d39a040d1762e04c4287a

Observation 27019421-34fd-42fc-86ab-9158bfcfd9bd · outbound

This paper cites AgentBench: Evaluating LLMs as agents.

Measuring Safety Alignment Effects in Autonomous Security Agents AgentBench: Evaluating LLMs as agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.354803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3fc0d58672a525007afe1620b0419e165f31365668b32600172dce9eda5b4804

Observation d3b29e2f-9f92-4800-9812-56596e3d0007 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

Measuring Safety Alignment Effects in Autonomous Security Agents ReAct: Synergizing reasoning and acting in language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.359667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3b94973e7b979d7f7a668998b80c7fbaf884a4d78f9aabf2c815e59d85821f9d

Observation 4dfd1d19-94b7-474f-8c4a-e53b2fcb1ece · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Measuring Safety Alignment Effects in Autonomous Security Agents Toolformer: Language models can teach themselves to use tools

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.364045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:74296a09152c5fe996b84ed21919ccf79a30f70e15a69bb477b764999ed1266c

Observation a8118503-132f-49ff-905c-f419a7d1f340 · outbound

This paper cites WebShop: Towards scalable real-world web interaction with grounded language agents.

Measuring Safety Alignment Effects in Autonomous Security Agents WebShop: Towards scalable real-world web interaction with grounded language agents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.336725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:9b46e127caa47c9a0a6e5f95438410f74d58ac181c867a52628bb77856e50ea6

Observation bfb5362c-a5df-471a-9d0b-4dc59079293f · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Measuring Safety Alignment Effects in Autonomous Security Agents Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.345780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:c58ebcbeceaab758f7775f1989bcf842c86bb77ba5a3be528dcfc90c5a9da8a6

Observation 661f81f5-df7c-42f5-8cd3-2737a4dc60ba · outbound

This paper cites InterCode: Standard- izing and benchmarking interactive coding with execution feedback.

Measuring Safety Alignment Effects in Autonomous Security Agents InterCode: Standard- izing and benchmarking interactive coding with execution feedback

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.311554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:9f12ee8e966e35ab612b633879a9511cdb9e3b9bcf9da2f1308a5c528f70d637

Observation b25b258d-9600-43bd-bc1f-17ddeab392cd · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.306636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:d97429c46746ca387ee074052a5b18c5febf5ae46357d1f7769ac1064a0903e2

Observation 4e592cfa-8a6e-427b-aab5-108aa9f137f0 · outbound

This paper cites Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press.

Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.316366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:844fbcc393610b85d2ffc09cc802a28ef5abf7518be809b2865eccfc3eb8298e

Observation 1379ec2b-efb3-4963-b406-5bd9979eb19a · outbound

This paper cites Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu.

Measuring Safety Alignment Effects in Autonomous Security Agents Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.350528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:36420824973e425c49172e80c25ae1613d1549ed7fbdf6f7f104befc396c83fc

Observation 2a5617d7-1b01-49d4-84b9-6b9a2a614570 · outbound

This paper cites AI Agents That Matter.

Measuring Safety Alignment Effects in Autonomous Security Agents AI Agents That Matter

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.909030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:210bc9f8bad8799b874053b86158a1c0de91e10ac09c9c904ce2eca40a300e50

Observation a0cd6025-45f8-44dc-9a09-b2e67efb4328 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Measuring Safety Alignment Effects in Autonomous Security Agents Refusal in language models is mediated by a single direction

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.388192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:6c5a3fdeda4f9b11bc33f90efa007a40566ecec69f0cc022fd235274ce7d3bcc

Observation f1545e90-6371-4de2-a540-e08bc201e08e · outbound

This paper cites Gemma 4: Byte for byte, the most capable open mod- els.

Measuring Safety Alignment Effects in Autonomous Security Agents Gemma 4: Byte for byte, the most capable open mod- els

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.292643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:5ba218cd41bfbe940ba18c9bd242f5b94c6107660dfe665201d639e9e090a4ee

Observation 2de2b40f-a39c-4519-b79a-c3eadddd3639 · outbound

This paper cites google/gemma-4-31b-it.

Measuring Safety Alignment Effects in Autonomous Security Agents google/gemma-4-31b-it

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.296953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:e3d72e2ae63940399230d043dead33555643af8b583e3ec522ffc2fbfc2ae5ee

Observation 3bbb32eb-4182-49a3-8793-914da6920925 · outbound

This paper cites google/gemma-4-26b-a4b-it.

Measuring Safety Alignment Effects in Autonomous Security Agents google/gemma-4-26b-a4b-it

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.301497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:1f351b97378e16e1683adbd112e27af00ac49ac6a792f54419ecbce6691e5d18

Observation 3a4d70e4-c6d5-4d12-b1fa-3d08fd5ba7a3 · outbound

This paper cites Gemma 4 uncensored.

Measuring Safety Alignment Effects in Autonomous Security Agents Gemma 4 uncensored

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.392340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:574022916edd762c2a5e7001632692bc3e03aa93c42dd6f652611ea4caf3d39a

Observation 85347c1c-327a-4c5e-be3a-73aaf70318da · outbound

This paper cites Trevorjs/gemma-4-31b-it-uncensored.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-31b-it-uncensored

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:33:23.312689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:ac327b18e364cbd3ee26527f2efc2feddbc1db9448cf71c2e10a00c1d67259bc

Observation b0b0aa9f-6419-46f4-b3c9-dff5906c0f21 · outbound

This paper cites Trevorjs/gemma-4-26b-a4b-it-uncensored.https://huggingface.co/TrevorJS/ gemma-4-26B-A4B-it-uncensored.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-26b-a4b-it-uncensored.https://huggingface.co/TrevorJS/ gemma-4-26B-A4B-it-uncensored

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.283158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:50a4c96b288e80ea2801a1a6f921dbae9bb40432bac816450d02ccda00a82b15

Observation f2cc92d1-e814-4abc-b757-27b5bfc4e98a · outbound

This paper cites unsloth/gemma-4-26b-a4b-it-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents unsloth/gemma-4-26b-a4b-it-gguf

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.287235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3e6dcec6aed5c7c0ab1e21231c0a7ebdb88fc091ce567329ddccacf452bcf82a

Observation 03a3de1a-533f-4214-99fb-2bbfdb4b912e · outbound

This paper cites Trevorjs/gemma-4-26b-a4b-it-uncensored-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-26b-a4b-it-uncensored-gguf

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.274516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:29368b924aa3a7024d763f2ad64770f011291cf581efafe881456a972b80242e

Observation 32258fff-5d6c-4ef5-adf5-fec5a139efe4 · outbound

This paper cites Qwen/qwen2.5-coder-7b-instruct-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents Qwen/qwen2.5-coder-7b-instruct-gguf

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.278713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:0bd2e5597933dfb0b5332b63f7c8e29bddb02d1cc22f3d5318d10e17df0d3074

Observation c9f19cfb-eeee-4a94-b47e-adc614e8c540 · outbound

This paper cites bartowski/qwen2.5-coder-7b-instruct-abliterated-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/qwen2.5-coder-7b-instruct-abliterated-gguf

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.269782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:a85d3a5475ae5b172cc83dc2e89cc3d0fdea28f9d9216abdf3a928bb425c2282

Observation fc0af708-fba0-4e45-bec9-b62df1096b9f · outbound

This paper cites bartowski/meta-llama-3.1-8b-instruct-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/meta-llama-3.1-8b-instruct-gguf

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.260335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:37f642c4289df64ebbdc5fe447903fe4b577ef71763defb187b278f559712edc

Observation ac746392-720e-46d4-8ae3-6314d678a31d · outbound

This paper cites bartowski/meta-llama-3.1-8b-instruct-abliterated-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/meta-llama-3.1-8b-instruct-abliterated-gguf

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.251233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:b5580c1a4b5df405aa6c06e4bc19df17d4d8570fc013260f300dc3a3aefdd6d3

Observation 237585d9-e280-4f38-9eaf-40615192bbf1 · outbound

This paper cites Trevorjs/gemma-4-31b-it-uncensored-gguf.

Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-31b-it-uncensored-gguf

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.255755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:f0345eeefc195687e74e5cbdbdac888b9ecdb6502419206815e1cd5918544070

Observation a4bc177a-d148-48e1-bb31-53b4ea71c81c · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Measuring Safety Alignment Effects in Autonomous Security Agents Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.888463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:5aa03774bfdbe7303694457342f9dae1ac903c8c72b6ee95f0922e2c70095d52

Observation 5536f233-0c18-4819-8a80-64ee0d9fa500 · outbound

This paper cites The GEM benchmark: Natural language generation, its evaluation and metrics.

Measuring Safety Alignment Effects in Autonomous Security Agents The GEM benchmark: Natural language generation, its evaluation and metrics

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.246895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:e96ebb0cfc05be943d41096179e0a17310d29ee974aa0773d7fc13a3f76d9874

Observation c4e7561d-cf84-46dd-b731-288d3f0d31d2 · outbound

This paper cites Inspect AI: Framework for large language model evaluations.

Measuring Safety Alignment Effects in Autonomous Security Agents Inspect AI: Framework for large language model evaluations

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.264756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:8d90438b658f16ae2147651cbaf762cc096d3a707c2259801b9236edd36554b7

Observation 0672b9c5-6fb0-439c-9e09-65e98af93209 · outbound

This paper cites Evaluating Frontier Models for Dangerous Capabilities.

Measuring Safety Alignment Effects in Autonomous Security Agents Evaluating Frontier Models for Dangerous Capabilities

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.916500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:68c294f698ac7a52be27472bb2c95d680c909fee62f51be560b191fe6e91f0ce

Observation ce90f471-643a-4f7f-a4ff-9105b12525d2 · outbound

This paper cites Model evaluation for extreme risks.

Measuring Safety Alignment Effects in Autonomous Security Agents Model evaluation for extreme risks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.922999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3f5672508ef42159b1ddeca1f9d25b3699024ca0fa944de6d6e0745ce0e06667

Observation 529e7e09-8a84-4916-b552-9c5359f675af · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Measuring Safety Alignment Effects in Autonomous Security Agents On the Opportunities and Risks of Foundation Models

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T04:28:05.901772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:7b892a9912085405125179b45ec4b423d2cf81693ab29eb549c4132e3a1b9d63

Observation 8d4e8e4e-0b82-49b3-913e-fe21b06442b2 · outbound

This paper cites Release Strategies and the Social Impacts of Language Models.

Measuring Safety Alignment Effects in Autonomous Security Agents Release Strategies and the Social Impacts of Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-20T04:28:05.929250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:b1278fd15befed5857f54fcefa58bf36a6f8d3a9876dded6525c4393a4d1f328

Observation 23f97ab3-c3d4-405e-9b4e-e7eda5c8e8a3 · outbound

This paper cites Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools.

Measuring Safety Alignment Effects in Autonomous Security Agents Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.895469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3cfc13ac1cdf18c9ecdbb12a690b4f070fa78608b89640efca074798ad4268a3

Observation 98654d1c-29cb-4718-a0a2-97625e5f0f9e · outbound

This paper cites Common weakness enumeration.

Measuring Safety Alignment Effects in Autonomous Security Agents Common weakness enumeration

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.242343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:a3134a581fb63ba2294018a55f498657f4697a5e879720d5804cc9a115315f87

Observation fcff3578-61f0-4486-98cb-12060f569178 · outbound

This paper cites OWASP Top 10: The ten most critical web application security risks.

Measuring Safety Alignment Effects in Autonomous Security Agents OWASP Top 10: The ten most critical web application security risks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.228976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:3c0bf03f5da17ee42b3ecbcde08e48a37a487bf4cd93e7f577a0e2172d632994

Observation 926c670f-a1fa-40de-b8e9-c4efa0b9b39d · outbound

This paper cites llama.cpp: LLM inference in C/C++.

Measuring Safety Alignment Effects in Autonomous Security Agents llama.cpp: LLM inference in C/C++

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.233038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:56ec36a38f591542f31118c348cbf1308ec75a818cd843d2f5c0731f7caf9df4

Observation f8821e5b-8e86-492d-9bbc-6d133934d26f · outbound

This paper cites The Hugging Face Hub: Machine learning collaboration platform.

Measuring Safety Alignment Effects in Autonomous Security Agents The Hugging Face Hub: Machine learning collaboration platform

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T04:28:07.237863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:7c7938f31525415abc6191fc8ffb308045bf8d1a5e8eb6822da34cdcc1b80195

Observation c9e4f9ae-021d-4d91-a3bc-c29a28ddec44 · outbound

This paper cites JSON Schema draft 2020-12.

Measuring Safety Alignment Effects in Autonomous Security Agents JSON Schema draft 2020-12

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-05-20T04:28:07.224764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:f7c219ef66e11bcfbb4f8fd09156b2d9a95feb042d4037eb3e239c836088cdd5

Pith citing papers

Observation f050463d-cd1d-4e14-9d4e-7d3f342d1085 · inbound

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs cites this paper.

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Measuring Safety Alignment Effects in Autonomous Security Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:22.293704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:25:22.293704Z digest=sha256:29a44bb98c9ffbcb6e7bcf16f5f146ee4cfc471be8cf2ae415cf277f78ebe4a2