Pith. sign in

Paper Citation Record · LEDGER

Geometry-Guided Constraint Learning for LLM Safety Classification

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.19366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19366 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:57:53.769079Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77878a51-adfe-4ce1-a2f4-87cffac742b1 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Geometry-Guided Constraint Learning for LLM Safety Classification Constitutional AI: Harmlessness from AI Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.367172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.367172Z digest=sha256:1e7dd4838fab1e5769532f4eb7721dc73f4648ea1ba834b8c58c359b1ece3464

Observation bd93902b-13eb-4f73-9cae-c70ede09db63 · outbound

This paper cites Dhillon, Joydeep Ghosh, and Suvrit Sra.

Geometry-Guided Constraint Learning for LLM Safety Classification Dhillon, Joydeep Ghosh, and Suvrit Sra

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.468624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.468624Z digest=sha256:badbd2f35934e8c6e9b93ee8e01c865527fad95d01506902972a05806d21eff1

Observation 4f2702e4-26be-4eb3-b53a-28150e736412 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023.

Geometry-Guided Constraint Learning for LLM Safety Classification Towards monosemanticity: Decomposing language models with dictionary learning.Transformer Circuits Thread, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.610755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.610755Z digest=sha256:d47db72143efe8b3888e268c733a8867d066e95a73b6811822e066690873c293

Observation 150d54bc-4265-44b9-862e-d8a5575a3f11 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Geometry-Guided Constraint Learning for LLM Safety Classification Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.775694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.775694Z digest=sha256:704110350773b2dac4ef1ecbaca2ef30565d4e81f1bd73d8d9a7902cc193e209

Observation 9d400a3f-0fe4-42db-8206-794618102f72 · outbound

This paper cites Learning Safety Constraints for Large Language Models.

Geometry-Guided Constraint Learning for LLM Safety Classification Learning Safety Constraints for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.928730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.928730Z digest=sha256:b202f7f54ee5a5fb71c443b99983e13b3be958b4cbb732919c6e438739d7f785

Observation 6e9adfc7-38c7-407d-ae5b-18d37ae03f35 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Geometry-Guided Constraint Learning for LLM Safety Classification Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.065478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.065478Z digest=sha256:42e5c858d1ab0a82a8cb384d38b879faacc04fd950c58670e66c7932ccfc74ff

Observation bb189223-be72-455a-9256-a01c19b0770e · outbound

This paper cites Toy models of superposition.Transformer Circuits Thread, 2022.

Geometry-Guided Constraint Learning for LLM Safety Classification Toy models of superposition.Transformer Circuits Thread, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.145415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.145415Z digest=sha256:46bd93d4439c2c6cd6cf350e46615e744185e2a1e98951e98b4505c54e94b64a

Observation 74081f2b-cef8-45c7-914c-27c9431aae08 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Geometry-Guided Constraint Learning for LLM Safety Classification WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.223916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.223916Z digest=sha256:ec1c148f9ade3e5c4518c246ae67d95e936a9bada1b132cb08b32fb44ca411f0

Observation f722afdd-539e-4cdb-8252-557146f93424 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Geometry-Guided Constraint Learning for LLM Safety Classification Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.312889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.312889Z digest=sha256:23c3542dd8d89bd359bc36aa4b0f47ae83e04bd63eb31efec11597fae703f76c

Observation 58c61947-c126-4803-a8eb-8ddab689f2e2 · outbound

This paper cites BeaverTails: Towards improved safety alignment of LLM via a human-preference dataset.

Geometry-Guided Constraint Learning for LLM Safety Classification BeaverTails: Towards improved safety alignment of LLM via a human-preference dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.439436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.439436Z digest=sha256:c9d628134cfad10a1b795b95cbcf358cd8bf5837db2bc64d06a5d2a24fe9876d

Observation 5d0d0f3a-666e-4ce8-be28-424ded230a3f · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.NeurIPS, 2024.

Geometry-Guided Constraint Learning for LLM Safety Classification Inference- time intervention: Eliciting truthful answers from a language model.NeurIPS, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.559072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.559072Z digest=sha256:7321cb6e648aae8cfa7ae4ea8ffa09e845905b2fb75a1010c2e7c5f3fc74fc83

Observation 0c8e92cc-e22c-4ed0-a184-d58a99a65790 · outbound

This paper cites A holistic approach to undesired content detection in the real world.AAAI, 2023.

Geometry-Guided Constraint Learning for LLM Safety Classification A holistic approach to undesired content detection in the real world.AAAI, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.647395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.647395Z digest=sha256:77b372ee3d7de176daf2293c8d9f54d552021692719108cae70117796ab49141

Observation a750113e-2e70-473d-9980-9ea70180cc5f · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Geometry-Guided Constraint Learning for LLM Safety Classification HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.737493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.737493Z digest=sha256:3db86d7751266dedd4a0707e24d91785acc0242271c4b234ef9ced1aed45e1cf

Observation f3671bb1-a417-463a-ad44-4b3f8319f5e1 · outbound

This paper cites Emergent Linear Representations in World Models of Self-Supervised Sequence Models.

Geometry-Guided Constraint Learning for LLM Safety Classification Emergent Linear Representations in World Models of Self-Supervised Sequence Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.868313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.868313Z digest=sha256:f29213713e85418eff479f7bf95977bfdc4365e7bd36075ff79ff8892f7e6e4b

Observation 7488ea7b-05ca-4470-9376-817762b3a2ba · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al.

Geometry-Guided Constraint Learning for LLM Safety Classification Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:52.958633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:52.958633Z digest=sha256:50539fbd8f41732f5799200da25d9734e9a17793d7687b48cedae88975bb5a45

Observation 1603682f-82ab-4b0a-9273-83061eca3aa8 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

Geometry-Guided Constraint Learning for LLM Safety Classification The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.017890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.017890Z digest=sha256:4b3ec32007fe1be338618ab5b0a64ddb05a27d961cccb9392f1f2f968cb68b45

Observation b0d48bd6-e79f-4ea3-93c0-7ebe0c72afd0 · outbound

This paper cites Qwen3 Technical Report.

Geometry-Guided Constraint Learning for LLM Safety Classification Qwen3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.086161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.086161Z digest=sha256:23257933693571e4a7374bf2a12bcad0e0292605341998cbd7f5a0281c47839f

Observation ab4c7147-9760-4409-b01c-cb4f85453004 · outbound

This paper cites NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails.

Geometry-Guided Constraint Learning for LLM Safety Classification NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.165247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.165247Z digest=sha256:fb0c2e4c20d52cb7eaa026dd51a7cdd9cbdf02405b804525022143446292bf40

Observation 95c63058-a6d1-4ec1-88ab-4b7763ef1908 · outbound

This paper cites Steering Language Models With Activation Engineering.

Geometry-Guided Constraint Learning for LLM Safety Classification Steering Language Models With Activation Engineering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.206211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.206211Z digest=sha256:51e1c97145ecdbdda123ab61deb7489272f594627e4579c710f1ee39d3228db6

Observation 944f9811-02ef-4b9c-bd97-9eaacec4baaa · outbound

This paper cites Jailbroken: How does LLM safety training fail?NeurIPS, 2024.

Geometry-Guided Constraint Learning for LLM Safety Classification Jailbroken: How does LLM safety training fail?NeurIPS, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.340057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.340057Z digest=sha256:d433028cf3d22063af034b7c83cf404cc83085eb013a1ba1b1fa78d17690a4ce

Observation c406c5c3-c19a-4e0c-ae8e-5de522131c7a · outbound

This paper cites Qwen2 Technical Report.

Geometry-Guided Constraint Learning for LLM Safety Classification Qwen2 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.490496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.490496Z digest=sha256:1e8917c79a2d0b898813fbe3bcb11e7be909c1d3abfc05eca2c96433f41c092e

Observation cd89a195-e45f-4813-88c7-9dbb2266b53c · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Geometry-Guided Constraint Learning for LLM Safety Classification Representation Engineering: A Top-Down Approach to AI Transparency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.641472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.641472Z digest=sha256:96cc80c9ac278bf79f440b85195dba0d4841ecfe11282a67adf04e9dcdb4a658

Observation 130d19d9-d915-4bd6-83f4-d7f57e8bdee1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Geometry-Guided Constraint Learning for LLM Safety Classification Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:53.769079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:53.769079Z digest=sha256:d52030af649488732bf9a496b146cdd6f40b0daf292a07580c0ee338e4af4e58

Pith citing papers

No inbound Pith citation observations are available.