Pith. sign in

Paper Citation Record · LEDGER

Adaptively Robust LLM Monitoring via Activation Watermarking

As of 8 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2603.23171.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.23171 v3

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T17:39:44.013978Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T05:57:19.399363Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T05:59:48.908825Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80302a1d-188c-4a7d-b4a5-ca56e33b47e1 · outbound

This paper cites GPT-4 Technical Report.

Adaptively Robust LLM Monitoring via Activation Watermarking GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.154929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.154929Z digest=sha256:ede88580b4a63a242d3ae25f100dc17e5e9c0bfc12957e26e9dd79d4423afd63

Observation 4d129cdd-2b53-43af-909f-e0d256dd8e59 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Adaptively Robust LLM Monitoring via Activation Watermarking Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.568606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.568606Z digest=sha256:f2b2b79e7ccc3c3cfbc0ca0cf0fef745f605c5c578f2d5d45b4c3a70d0ef5395

Observation 168ce311-8d49-4ef1-9c82-cc66d1aa94dc · outbound

This paper cites MirrorCheck: Efficient Adversarial Defense for Vision-Language Models.

Adaptively Robust LLM Monitoring via Activation Watermarking MirrorCheck: Efficient Adversarial Defense for Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.825588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.825588Z digest=sha256:8bad1f4dfeb3f0e897d49cc0746531ab250940073b4d8e53f33048e870aefeb8

Observation 2130c3f7-ec79-428c-a265-1d11bb15dfc9 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Adaptively Robust LLM Monitoring via Activation Watermarking Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.949795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.949795Z digest=sha256:ec026407dc9926f9ebe36b60090dfb50ffa60ff9f6e9cda330658ec2b2445a8a

Observation c1a959fb-3477-4dfc-a69c-75e4b35d5cf3 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Adaptively Robust LLM Monitoring via Activation Watermarking Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.084298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.084298Z digest=sha256:78bbb86e0ff424a2397cd381fb9d5084243e6dfe3c88b9c8a4735503b5709219

Observation 2e4e9bf4-0386-432b-8f8c-e4ab437f8c22 · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

Adaptively Robust LLM Monitoring via Activation Watermarking HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.343655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.343655Z digest=sha256:f524db8815cc30bd7d490d91d0ddc3a53a820d9eb28f82f842d2823ed0104f65

Observation 6aafcc01-92ce-4356-a137-8f3106c88bbb · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Adaptively Robust LLM Monitoring via Activation Watermarking DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.537821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.537821Z digest=sha256:8d7d6c76978c2ae9dfbf9d385a7b33588d3419020a2dbf285759816dccc10a73

Observation cdd5f30b-c4b5-425a-ac88-4e8af75bfd13 · outbound

This paper cites Against The Achilles' Heel: A Survey on Red Teaming for Generative Models.

Adaptively Robust LLM Monitoring via Activation Watermarking Against The Achilles' Heel: A Survey on Red Teaming for Generative Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.714332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.714332Z digest=sha256:698b9f5f6fe7323086b13a40addbf73c94b583c373a7ed121b32b586621a4e3f

Observation 4bd3435b-6a83-4460-a7b4-e1128aa18ced · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Adaptively Robust LLM Monitoring via Activation Watermarking AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.780056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.780056Z digest=sha256:7f83d69cb840b42389648fd168d456bda10b584ab826c5688556c53ab31c1325

Observation 2cdf4dad-04a2-40a0-9eef-ad3ff45a8632 · outbound

This paper cites The Llama 3 Herd of Models.

Adaptively Robust LLM Monitoring via Activation Watermarking The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.876621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.876621Z digest=sha256:d7cf7c1cd86ccaf7032ecaf394c4f0ccca1e2b480d495c86fa381f3b927912d0

Observation 443d7cf7-31b2-4a71-9d84-65be7c98f303 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Adaptively Robust LLM Monitoring via Activation Watermarking XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.019845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.019845Z digest=sha256:32cb4c42d3845f8efb02f56c989062808f8a30fe3771ede0d5501335676a2767

Observation f7f84e7b-ee6e-4ae1-ad62-eb45eb2fefd5 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

Adaptively Robust LLM Monitoring via Activation Watermarking Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.163305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.163305Z digest=sha256:03abbabd9aa4cff4f732a8a93f1580faf9c5f3f14730f752237b02ad8c7f5352

Observation e8786eb1-1888-4eff-be47-7eddf6def933 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Adaptively Robust LLM Monitoring via Activation Watermarking Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.299576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.299576Z digest=sha256:73b73dadbfe70d8fdf6f368ab4c1ad28346367fbc1e09fb388d6d1dd685018ff

Observation 80efb0bd-9825-4256-92e4-a6d350c0c80c · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Adaptively Robust LLM Monitoring via Activation Watermarking MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.379650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.379650Z digest=sha256:a545a47991ea825c2e8ae2977a53c6c7ecc43eea3225032662b4f3dce010675f

Observation 089d5599-51c5-4e85-863e-53799a233b26 · outbound

This paper cites Qwen3Guard Technical Report.

Adaptively Robust LLM Monitoring via Activation Watermarking Qwen3Guard Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.471239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.471239Z digest=sha256:62a403eefc8244189cda3cdf24bf883f7b293fe32ae8368be71feacd3b40b6ba

Observation 25cfcb93-92fa-4f6b-849e-c6e6a8d85b96 · outbound

This paper cites SoK: Watermarking for AI-Generated Content.

Adaptively Robust LLM Monitoring via Activation Watermarking SoK: Watermarking for AI-Generated Content

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.550309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.550309Z digest=sha256:f3b5f34d5c1589e6a3616b245d4ab742dfdadf93b6dbd04e0a7410bb421f430d

Observation 9d3951d5-eaf9-4f24-ab2a-e4036b427202 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Adaptively Robust LLM Monitoring via Activation Watermarking Instruction-Following Evaluation for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.658628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.658628Z digest=sha256:5739b61bedfcfaacd1edcf6a61690a067d2166b31f78abb880a5219534a04e1d

Observation c1e5f8ce-f3d7-45ee-be63-6b742e14bc9d · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Adaptively Robust LLM Monitoring via Activation Watermarking Representation Engineering: A Top-Down Approach to AI Transparency

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.766293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.766293Z digest=sha256:8aeefd240982449d78f456b77961dc00ac0652b451856fda1f84303d3b868097

Observation b4f38a60-5ee7-47a6-a6f4-c4d7d2608618 · outbound

This paper cites an unresolved cited work.

Adaptively Robust LLM Monitoring via Activation Watermarking Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.884154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.884154Z digest=sha256:0ec2a96ca76c40729e62f78464a67bfb9ab52711f2ba2c9bf113ca77597f4fdf

Observation c05d3460-5916-4c79-a27e-666356ee0d6b · outbound

This paper cites The attacker then issues the translated prompt to the model and, for evaluation, translates the answer back into English.

Adaptively Robust LLM Monitoring via Activation Watermarking The attacker then issues the translated prompt to the model and, for evaluation, translates the answer back into English

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:44.013978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:44.013978Z digest=sha256:f0bfb690416da6b283e9887fbc657342f3129cf7c0404199d088ef5e4c357246

Observation 816375f6-b7aa-484b-9df8-095c7c70de76 · outbound

This paper cites Optimizing Adaptive Attacks against Watermarks for Language Models.

Adaptively Robust LLM Monitoring via Activation Watermarking Optimizing Adaptive Attacks against Watermarks for Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.676660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.676660Z digest=sha256:d8d37db0b6ebcb4043651f5626791a7cc19239fa3d0d5e618224f6379987ecea

Observation 24a9af66-5191-471a-93b5-341c475e8be1 · outbound

This paper cites Obfuscated Activations Bypass LLM Latent-Space Defenses.

Adaptively Robust LLM Monitoring via Activation Watermarking Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.404497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.404497Z digest=sha256:0484ea6ac179a3f57476e1353f7ed70a79439b585baeb8ef3f1b02c7d854de92

Observation 1da578b5-3c97-45b9-a11b-dcc30e0382f2 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.arXiv preprint arXiv:2307.04657,.

Adaptively Robust LLM Monitoring via Activation Watermarking Beavertails: Towards improved safety alignment of llm via a human-preference dataset.arXiv preprint arXiv:2307.04657,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.200858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.200858Z digest=sha256:0716d3bdcecc7df8fa663606af59a14121770350d2809f799725d8021315727d

Observation 9303b4d7-e24e-4ab1-9ff5-947e0386b362 · outbound

This paper cites Mitigating Watermark Forgery in Generative Models via Randomized Key Selection.

Adaptively Robust LLM Monitoring via Activation Watermarking Mitigating Watermark Forgery in Generative Models via Randomized Key Selection

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.297826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.297826Z digest=sha256:541de24bcef3ef939d9c2e003f00b711171f6b72c8aba6f016223b601f477d86

Observation 0a045218-3b2d-4b1f-a4e2-4c2ded9bac89 · outbound

This paper cites ISBN 9798400720406.

Adaptively Robust LLM Monitoring via Activation Watermarking ISBN 9798400720406

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.434758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.434758Z digest=sha256:06b6e73fc4d9f0f617c4c4e524477036fa6fd82090c1f885b77fee2f49301dee

Pith citing papers

Observation a946a793-65b3-4915-9170-80454a373aa6 · inbound

Watermarking Should Be Treated as a Monitoring Primitive cites this paper.

Watermarking Should Be Treated as a Monitoring Primitive Adaptively Robust LLM Monitoring via Activation Watermarking

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-30T02:04:38.827270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:10:59.937855Z digest=sha256:27d75b8a0485e5536b4f8aeb09086899e49acf625c6e3eb205d296c51cec73e7

Observation 54f1da3f-0259-4820-aa44-f130864553fe · inbound

Watermarking Should Be Treated as a Monitoring Primitive cites this paper.

Watermarking Should Be Treated as a Monitoring Primitive Adaptively Robust LLM Monitoring via Activation Watermarking

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-30T02:04:38.827270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:57:19.399363Z digest=sha256:d36ec884c0bedb6b312251015a4696ec498114fdabc409f12d93a4608607f047