Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T17:39:44.013978Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2603.23171.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T17:39:44.013978Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-15T05:57:19.399363Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T05:59:48.908825Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80302a1d-188c-4a7d-b4a5-ca56e33b47e1 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d129cdd-2b53-43af-909f-e0d256dd8e59 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168ce311-8d49-4ef1-9c82-cc66d1aa94dc · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2130c3f7-ec79-428c-a265-1d11bb15dfc9 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Measuring Mathematical Problem Solving With the MATH Dataset
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1a959fb-3477-4dfc-a69c-75e4b35d5cf3 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e4e9bf4-0386-432b-8f8c-e4ab437f8c22 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aafcc01-92ce-4356-a137-8f3106c88bbb · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd5f30b-c4b5-425a-ac88-4e8af75bfd13 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd3435b-6a83-4460-a7b4-e1128aa18ced · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cdf4dad-04a2-40a0-9eef-ad3ff45a8632 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 443d7cf7-31b2-4a71-9d84-65be7c98f303 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f84e7b-ee6e-4ae1-ad62-eb45eb2fefd5 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8786eb1-1888-4eff-be47-7eddf6def933 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80efb0bd-9825-4256-92e4-a6d350c0c80c · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089d5599-51c5-4e85-863e-53799a233b26 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Qwen3Guard Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25cfcb93-92fa-4f6b-849e-c6e6a8d85b96 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking SoK: Watermarking for AI-Generated Content
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d3951d5-eaf9-4f24-ab2a-e4036b427202 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Instruction-Following Evaluation for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e5f8ce-f3d7-45ee-be63-6b742e14bc9d · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Representation Engineering: A Top-Down Approach to AI Transparency
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f38a60-5ee7-47a6-a6f4-c4d7d2608618 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05d3460-5916-4c79-a27e-666356ee0d6b · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking The attacker then issues the translated prompt to the model and, for evaluation, translates the answer back into English
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 816375f6-b7aa-484b-9df8-095c7c70de76 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Optimizing Adaptive Attacks against Watermarks for Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24a9af66-5191-471a-93b5-341c475e8be1 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Obfuscated Activations Bypass LLM Latent-Space Defenses
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da578b5-3c97-45b9-a11b-dcc30e0382f2 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Beavertails: Towards improved safety alignment of llm via a human-preference dataset.arXiv preprint arXiv:2307.04657,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9303b4d7-e24e-4ab1-9ff5-947e0386b362 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a045218-3b2d-4b1f-a4e2-4c2ded9bac89 · outbound
Adaptively Robust LLM Monitoring via Activation Watermarking ISBN 9798400720406
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a946a793-65b3-4915-9170-80454a373aa6 · inbound
Watermarking Should Be Treated as a Monitoring Primitive Adaptively Robust LLM Monitoring via Activation Watermarking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 54f1da3f-0259-4820-aa44-f130864553fe · inbound
Watermarking Should Be Treated as a Monitoring Primitive Adaptively Robust LLM Monitoring via Activation Watermarking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.