Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:52:47.572246Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2508.19980.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:52:47.572246Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f42fe708-dd30-4539-80ac-723e359d7cb4 · outbound
Evaluating Language Model Reasoning about Confidential Information Prompt leakage effect and mitigation strategies for multi-turn LLM ap- plications
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2e0e3fe2-f9c5-4427-930f-18edc4e2f26c · outbound
Evaluating Language Model Reasoning about Confidential Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6669c185-6eed-49d6-8c25-434a3bc8cafb · outbound
Evaluating Language Model Reasoning about Confidential Information Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1648d5bb-b74f-45b2-a6d9-bea32523d993 · outbound
Evaluating Language Model Reasoning about Confidential Information Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96214bf6-1ba4-4227-b11d-cb9418ba1322 · outbound
Evaluating Language Model Reasoning about Confidential Information DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f73a527-36f8-4945-900e-21f1b8fb4fc0 · outbound
Evaluating Language Model Reasoning about Confidential Information GPT-4o System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c758c7d7-f648-4164-b243-fede705e2b77 · outbound
Evaluating Language Model Reasoning about Confidential Information Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71410a8-08ee-4239-b6c5-5d7cd5c2f155 · outbound
Evaluating Language Model Reasoning about Confidential Information Safety pretraining: Toward the next generation of safe ai
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75a303ba-dd04-419d-bd9a-39f6c3661454 · outbound
Evaluating Language Model Reasoning about Confidential Information HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e913d6-27af-4af8-a544-46b7653a9ca4 · outbound
Evaluating Language Model Reasoning about Confidential Information Can LLMs Follow Simple Rules?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36dc7f6c-2c5f-4a6f-8fd9-cb615e8658c4 · outbound
Evaluating Language Model Reasoning about Confidential Information A Closer Look at System Prompt Robustness
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a7f098c-9bc3-4288-9c08-67444ce3c695 · outbound
Evaluating Language Model Reasoning about Confidential Information SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98454188-99f4-4c45-9a2c-f51f54524fb9 · outbound
Evaluating Language Model Reasoning about Confidential Information Jailbreaking LLM-Controlled Robots
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2795196e-45e6-458d-8120-f39063b2c735 · outbound
Evaluating Language Model Reasoning about Confidential Information Predicting the performance of black-box llms through self-queries
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c66cd3-aaeb-4702-9cf6-0ee28df97a34 · outbound
Evaluating Language Model Reasoning about Confidential Information Antidistillation sampling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c401e31-80fd-42d7-98a8-32896ec4d246 · outbound
Evaluating Language Model Reasoning about Confidential Information Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d29d1f-de95-49ea-acde-df0167319f14 · outbound
Evaluating Language Model Reasoning about Confidential Information Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35815867-dbe3-4cb1-b11d-5136d5061d30 · outbound
Evaluating Language Model Reasoning about Confidential Information Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f7fe58f-43e0-481c-8450-15a2d8b072f8 · outbound
Evaluating Language Model Reasoning about Confidential Information The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94dfdbb-0546-42e0-9f77-aca3bab65642 · outbound
Evaluating Language Model Reasoning about Confidential Information Qwen3 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90908afd-a11e-436f-8dff-d9fc765213f3 · outbound
Evaluating Language Model Reasoning about Confidential Information Trading Inference-Time Compute for Adversarial Robustness
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a7b0b7-37ba-47da-9ea0-588d5c74fecd · outbound
Evaluating Language Model Reasoning about Confidential Information Backtracking Improves Generation Safety
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb9024e3-ba0c-4e5d-b643-bb0ff2c21409 · outbound
Evaluating Language Model Reasoning about Confidential Information Instruction-Following Evaluation for Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1be5d31f-5cc8-4aea-94b0-27b1d288a608 · outbound
Evaluating Language Model Reasoning about Confidential Information Large Language Models can Learn Rules
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b5ed83-e030-443a-99be-4129d2ea2b33 · outbound
Evaluating Language Model Reasoning about Confidential Information Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add381a1-d232-4eea-96cb-6a2a6f52daf3 · outbound
Evaluating Language Model Reasoning about Confidential Information Rating: [[rating]]
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9963bfef-e6de-4133-819e-e0164ac93a29 · outbound
Evaluating Language Model Reasoning about Confidential Information However, we focus on cases where we do not have knowledge of the specialized target string
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dcb888b3-01ac-4a48-af50-5c066515b242 · outbound
Evaluating Language Model Reasoning about Confidential Information Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce381b0-ab56-4b71-8529-27ee8027dfce · outbound
Evaluating Language Model Reasoning about Confidential Information Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1595a515-9db7-48a2-bbe1-426554394451 · outbound
Evaluating Language Model Reasoning about Confidential Information JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572586c9-8a08-469f-8814-8ada343975f8 · outbound
Evaluating Language Model Reasoning about Confidential Information Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8400fb9-4462-45b1-9ed0-8d82afcc5b6f · outbound
Evaluating Language Model Reasoning about Confidential Information Stress-Testing Capability Elicitation With Password-Locked Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.