Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T11:57:53.769079Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.19366.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T11:57:53.769079Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77878a51-adfe-4ce1-a2f4-87cffac742b1 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Constitutional AI: Harmlessness from AI Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd93902b-13eb-4f73-9cae-c70ede09db63 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Dhillon, Joydeep Ghosh, and Suvrit Sra
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2702e4-26be-4eb3-b53a-28150e736412 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150d54bc-4265-44b9-862e-d8a5575a3f11 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d400a3f-0fe4-42db-8206-794618102f72 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Learning Safety Constraints for Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e9adfc7-38c7-407d-ae5b-18d37ae03f35 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb189223-be72-455a-9256-a01c19b0770e · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Toy models of superposition.Transformer Circuits Thread, 2022
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74081f2b-cef8-45c7-914c-27c9431aae08 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f722afdd-539e-4cdb-8252-557146f93424 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58c61947-c126-4803-a8eb-8ddab689f2e2 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification BeaverTails: Towards improved safety alignment of LLM via a human-preference dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d0d0f3a-666e-4ce8-be28-424ded230a3f · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Inference- time intervention: Eliciting truthful answers from a language model.NeurIPS, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8e92cc-e22c-4ed0-a184-d58a99a65790 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification A holistic approach to undesired content detection in the real world.AAAI, 2023
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a750113e-2e70-473d-9980-9ea70180cc5f · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3671bb1-a417-463a-ad44-4b3f8319f5e1 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Emergent Linear Representations in World Models of Self-Supervised Sequence Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7488ea7b-05ca-4470-9376-817762b3a2ba · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1603682f-82ab-4b0a-9273-83061eca3aa8 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification The Linear Representation Hypothesis and the Geometry of Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d48bd6-e79f-4ea3-93c0-7ebe0c72afd0 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Qwen3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab4c7147-9760-4409-b01c-cb4f85453004 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c63058-a6d1-4ec1-88ab-4b7763ef1908 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Steering Language Models With Activation Engineering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 944f9811-02ef-4b9c-bd97-9eaacec4baaa · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Jailbroken: How does LLM safety training fail?NeurIPS, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c406c5c3-c19a-4e0c-ae8e-5de522131c7a · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Qwen2 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd89a195-e45f-4813-88c7-9dbb2266b53c · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Representation Engineering: A Top-Down Approach to AI Transparency
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130d19d9-d915-4bd6-83f4-d7f57e8bdee1 · outbound
Geometry-Guided Constraint Learning for LLM Safety Classification Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.