Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:45.016767Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2502.07985.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:45.016767Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
17 of 17 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 2022d310-a28d-4ce0-8134-f1ba8f434780 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599b5898-dc2b-4146-843a-b5fcd295d9aa · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d282f1f9-1687-410f-9a03-37dda913f7c1 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7d7c88-ac9b-42c2-9fa1-c29ab2ee356b · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Fight Back Against Jailbreaking via Prompt Adversarial Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b5db248-9976-4374-b751-22a3d133c73a · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec7ad2e0-4ac6-4a6c-9c60-59326b626e3b · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models ” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 169c8b7a-e174-4a73-a2ab-c6fe5b378ac2 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d99bd8a6-60a0-4f3d-ae45-01b02d561c2a · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models The ART of LLM Refinement: Ask, Refine, and Trust
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2335f18c-ae00-4d3a-aa4c-0bb93ef66df7 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a488f1-08c4-43a6-a2a4-94b88f78a080 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f65bae9-4b97-4834-8339-ca1575789c22 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61923c79-eda0-4f66-ace6-2420611a4b4f · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models t spect 0 Safety and harmless
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d0a213c-c91b-4670-93f4-8d7e9cdff5f3 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Configurable Safety Tuning of Language Models with Synthetic Preference Data
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35a6b2f9-8d39-4f06-adff-8953b346837d · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87375eb2-47a2-4606-a742-5d48d49c85f9 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models A Survey on LLM-as-a-Judge
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc5c325-868a-42c1-87dc-236572a5aafb · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf4a044-44f1-432b-a217-65ca0d19f1f8 · outbound
MetaSC: Test-Time Safety Specification Optimization for Language Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd0c1e7-a6d4-41c3-87cf-042ecdf831cb · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models MetaSC: Test-Time Safety Specification Optimization for Language Models
Reference 194
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba2fea89-06d9-46f3-aa23-ece80465df79 · inbound
Prompt Governance? On Governing Technologies Governed by Natural Language MetaSC: Test-Time Safety Specification Optimization for Language Models
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.