Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T17:23:52.388431Z
Paper Citation Record · LEDGER
As of 23 July 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2606.01060.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T17:23:52.388431Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ecd05710-950f-4dab-ae8f-c74a8fd69606 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Locating and Editing Factual Associations in
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb246c78-34b2-48d6-87d5-7d5ea22e689a · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , pages =
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca7dd34-696f-4983-b581-e3521e117e04 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Alignment Quality Index ( AQI ) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c570b390-195c-455d-a83c-c9f7db448ab8 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Shai and Sarah E
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c509664a-987b-47ec-b0eb-42feb21884fc · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 54d52469-bcff-45c6-be6b-ec2b969b7c84 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 43e4c1d8-3df4-4160-90f9-6c05ef7ae6f5 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Refusal in Language Models Is Mediated by a Single Direction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation be735694-a87a-4112-b82c-a4ec9e68329f · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 79df4d10-a3f6-4352-a823-9515a66dbae6 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Representation Engineering: A Top-Down Approach to AI Transparency
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 49101334-9e51-4e7e-9a66-9295d9266fdc · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3cdc3cc0-8b62-4d2c-99a1-2323df1b044a · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Advances in Neural Information Processing Systems, Datasets and Benchmarks Track , year =
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ebaab9a-d133-4fdb-9ca5-8516878ab476 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models 2 OLMo 2 Furious
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5970e424-eaa2-48bc-b9ca-34279584c393 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Mistral 7B
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f8eea177-f48a-4510-8de8-7b901c7361f9 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d2cc64b3-9d15-44e6-956f-be034d003f9c · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c3da261a-bc1d-4072-b5f8-49e1a114b364 · outbound
MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Similarity of Neural Network Representations Revisited
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
No inbound Pith citation observations are available.