Pith. sign in

Paper Citation Record · LEDGER

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models

As of 23 July 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2606.01060.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01060 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:23:52.388431Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecd05710-950f-4dab-ae8f-c74a8fd69606 · outbound

This paper cites Locating and Editing Factual Associations in.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Locating and Editing Factual Associations in

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T17:23:52.388431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:d537fa5fa5490bcc39fe7b0e4fb0adac2f9b3125b5446e3a8480cd96aeb9ea43

Observation bb246c78-34b2-48d6-87d5-7d5ea22e689a · outbound

This paper cites Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , pages =.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , pages =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T17:23:52.388431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:3e75e847fbe51c7565750107d6d844e7329c10172877cb40218eafd28353ed91

Observation cca7dd34-696f-4983-b581-e3521e117e04 · outbound

This paper cites Alignment Quality Index ( AQI ) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Alignment Quality Index ( AQI ) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations

Reference 3

Resolution
verified exact
doi, observed 2026-06-28T17:32:25.095198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:f8a6139982208747e64bd242ad925bb790afcca8a64eca37c2957978939d72e9

Observation c570b390-195c-455d-a83c-c9f7db448ab8 · outbound

This paper cites Shai and Sarah E.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Shai and Sarah E

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T17:23:52.388431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:d677b2276814840dc355bab87fcbcce2402998a0922ab8124fae3441b05caabd

Observation c509664a-987b-47ec-b0eb-42feb21884fc · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.711820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:f1ca8281bcd3a164583894e84b4920ebe5fda2e8e1f7acd97fd66b5730131ce1

Observation 54d52469-bcff-45c6-be6b-ec2b969b7c84 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.704818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:fe9511c02de3afd73fd7ff2ade93633156db3e03fefa232a30477e9cc32d4784

Observation 43e4c1d8-3df4-4160-90f9-6c05ef7ae6f5 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Refusal in Language Models Is Mediated by a Single Direction

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:16:13.707605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:8c548a913cffc41b3f7467d329fbac08454a5032b15773acdb6f11bd5566a390

Observation be735694-a87a-4112-b82c-a4ec9e68329f · outbound

This paper cites Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.716883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:bf37f0cf7ddbccc4649701a12ad35d7475a784798e1953d94c346a6f8ebf378f

Observation 79df4d10-a3f6-4352-a823-9515a66dbae6 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:16:13.713853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:7bf102cd86c770cbd4baa343194cf27dcb0f13cdbaa2b81749b0f529bb1d9939

Observation 49101334-9e51-4e7e-9a66-9295d9266fdc · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:16:13.702515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:7106e4d7a3fd3c685c4a1bf6f481d60cc0f4a3dfb5cd07dd24396a960b4b6a72

Observation 3cdc3cc0-8b62-4d2c-99a1-2323df1b044a · outbound

This paper cites Advances in Neural Information Processing Systems, Datasets and Benchmarks Track , year =.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Advances in Neural Information Processing Systems, Datasets and Benchmarks Track , year =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T17:23:52.388431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:af3e79105b76dbb1ad5a962602202fb0e7233017787e104fb41fed5ef75bb45c

Observation 4ebaab9a-d133-4fdb-9ca5-8516878ab476 · outbound

This paper cites 2 OLMo 2 Furious.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models 2 OLMo 2 Furious

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:13.699367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:127387c937532a95afebc005ff68f88ac8546496e9d421d478ca607a287eaabb

Observation 5970e424-eaa2-48bc-b9ca-34279584c393 · outbound

This paper cites Mistral 7B.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Mistral 7B

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:16:13.714417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:5b8c7c80445407094face524c2863ec978316dbcfc3d75511610bc81dc893600

Observation f8eea177-f48a-4510-8de8-7b901c7361f9 · outbound

This paper cites The Llama 3 Herd of Models.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models The Llama 3 Herd of Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:16:13.707964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:8bcfa125b87265da773f4d6d008607986671e71da9b0f94e415a2d44121254dd

Observation d2cc64b3-9d15-44e6-956f-be034d003f9c · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:16:13.717202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:7a804751f8f42b0efdc398c7b72c553490b3721bd058741bd83c04d92b515169

Observation c3da261a-bc1d-4072-b5f8-49e1a114b364 · outbound

This paper cites Similarity of Neural Network Representations Revisited.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Similarity of Neural Network Representations Revisited

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:16:13.697037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:057fb6894f3f16172180ec833c79b3cc40aa723838b7f6a3d903fd5a24ce1df9

Pith citing papers

No inbound Pith citation observations are available.