Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T04:24:46.357505Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2605.19722.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T04:24:46.357505Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T22:25:22.293704Z
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 220f1528-4a7f-49d4-820b-3050b76235a8 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Christiano, Jan Leike, Tom B
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 418ad41f-d931-4834-b4f0-24ca7d35dc73 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Fine-Tuning Language Models from Human Preferences
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bba8e7d0-162c-4b34-8a43-9543a106f6d9 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3804f4c3-570d-400c-a2b2-aad6e6729709 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents A General Language Assistant as a Laboratory for Alignment
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f008e3e8-6c37-4891-828e-426bb8eabf12 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ec530f60-5bc1-457e-83ff-ee2c48e621fb · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9c3e0a6e-49b3-4230-8e44-b6d404b88b38 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents TruthfulQA: Measuring how models mimic human falsehoods
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 95ec4f96-4b7a-4caf-924c-c79700b78a70 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Aligning AI with shared human values
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cda6f5ab-4562-4c1c-80b1-6d8beda9763b · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Ethical and social risks of harm from Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f964082b-0be8-4452-835d-1d262726894a · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents XSTest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b9ad7247-074b-4fde-9dc7-fe6fd2c6d444 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents HarmBench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ebc3ab62-8b90-4211-b0dd-f746002d709b · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6e90c35a-427c-4e0b-a388-5e9e9fe3088f · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents A StrongREJECT for empty jailbreaks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 17527ca2-a382-41aa-b45d-139de78f1618 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2849437a-18b7-4d90-9859-d3229a6cc7b0 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems 36
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 33321e15-c800-4725-ace6-6b33edf20776 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1fd1169f-7319-49c8-9434-816def9d1576 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 41639c6d-07f6-4455-9fcc-daca9c01376f · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 76a31408-5d3c-4e96-98d8-fd9c9e15328a · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents NYU CTF bench: A scalable open-source benchmark dataset for evaluating LLMs in offensive security
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 400283d2-2649-4040-8004-be19f07e0161 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, Farshad Khorrami, Prashanth Krishnamurthy, Brendan Dolan-Gavitt, Muhammad Shafique, Karthik R
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f03c3be7-038e-4f54-a246-c534a2308f11 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents SEC-bench: Automated bench- marking of LLM agents on real-world software security tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 27019421-34fd-42fc-86ab-9158bfcfd9bd · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents AgentBench: Evaluating LLMs as agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d3b29e2f-9f92-4800-9812-56596e3d0007 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents ReAct: Synergizing reasoning and acting in language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4dfd1d19-94b7-474f-8c4a-e53b2fcb1ece · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Toolformer: Language models can teach themselves to use tools
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a8118503-132f-49ff-905c-f419a7d1f340 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents WebShop: Towards scalable real-world web interaction with grounded language agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bfb5362c-a5df-471a-9d0b-4dc59079293f · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 661f81f5-df7c-42f5-8cd3-2737a4dc60ba · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents InterCode: Standard- izing and benchmarking interactive coding with execution feedback
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b25b258d-9600-43bd-bc1f-17ddeab392cd · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4e592cfa-8a6e-427b-aab5-108aa9f137f0 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1379ec2b-efb3-4963-b406-5bd9979eb19a · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2a5617d7-1b01-49d4-84b9-6b9a2a614570 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents AI Agents That Matter
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a0cd6025-45f8-44dc-9a09-b2e67efb4328 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Refusal in language models is mediated by a single direction
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1545e90-6371-4de2-a540-e08bc201e08e · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Gemma 4: Byte for byte, the most capable open mod- els
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2de2b40f-a39c-4519-b79a-c3eadddd3639 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents google/gemma-4-31b-it
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3bbb32eb-4182-49a3-8793-914da6920925 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents google/gemma-4-26b-a4b-it
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a4d70e4-c6d5-4d12-b1fa-3d08fd5ba7a3 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Gemma 4 uncensored
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 85347c1c-327a-4c5e-be3a-73aaf70318da · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-31b-it-uncensored
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b0b0aa9f-6419-46f4-b3c9-dff5906c0f21 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-26b-a4b-it-uncensored.https://huggingface.co/TrevorJS/ gemma-4-26B-A4B-it-uncensored
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2cc92d1-e814-4abc-b757-27b5bfc4e98a · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents unsloth/gemma-4-26b-a4b-it-gguf
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 03a3de1a-533f-4214-99fb-2bbfdb4b912e · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-26b-a4b-it-uncensored-gguf
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 32258fff-5d6c-4ef5-adf5-fec5a139efe4 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Qwen/qwen2.5-coder-7b-instruct-gguf
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c9f19cfb-eeee-4a94-b47e-adc614e8c540 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/qwen2.5-coder-7b-instruct-abliterated-gguf
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fc0af708-fba0-4e45-bec9-b62df1096b9f · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/meta-llama-3.1-8b-instruct-gguf
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ac746392-720e-46d4-8ae3-6314d678a31d · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents bartowski/meta-llama-3.1-8b-instruct-abliterated-gguf
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 237585d9-e280-4f38-9eaf-40615192bbf1 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Trevorjs/gemma-4-31b-it-uncensored-gguf
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a4bc177a-d148-48e1-bb31-53b4ea71c81c · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5536f233-0c18-4819-8a80-64ee0d9fa500 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents The GEM benchmark: Natural language generation, its evaluation and metrics
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c4e7561d-cf84-46dd-b731-288d3f0d31d2 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Inspect AI: Framework for large language model evaluations
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0672b9c5-6fb0-439c-9e09-65e98af93209 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Evaluating Frontier Models for Dangerous Capabilities
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ce90f471-643a-4f7f-a4ff-9105b12525d2 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Model evaluation for extreme risks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 529e7e09-8a84-4916-b552-9c5359f675af · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents On the Opportunities and Risks of Foundation Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8d4e8e4e-0b82-49b3-913e-fe21b06442b2 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Release Strategies and the Social Impacts of Language Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23f97ab3-c3d4-405e-9b4e-e7eda5c8e8a3 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 98654d1c-29cb-4718-a0a2-97625e5f0f9e · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents Common weakness enumeration
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fcff3578-61f0-4486-98cb-12060f569178 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents OWASP Top 10: The ten most critical web application security risks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 926c670f-a1fa-40de-b8e9-c4efa0b9b39d · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents llama.cpp: LLM inference in C/C++
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f8821e5b-8e86-492d-9bbc-6d133934d26f · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents The Hugging Face Hub: Machine learning collaboration platform
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c9e4f9ae-021d-4d91-a3bc-c29a28ddec44 · outbound
Measuring Safety Alignment Effects in Autonomous Security Agents JSON Schema draft 2020-12
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f050463d-cd1d-4e14-9d4e-7d3f342d1085 · inbound
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs Measuring Safety Alignment Effects in Autonomous Security Agents
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.