Pith. sign in

Paper Citation Record · LEDGER

Towards eliciting latent knowledge from LLMs with mechanistic interpretability

As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 9 inbound Pith citation observations for arXiv:2505.14352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14352 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:40:07.861843Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:20:51.989267Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1f0d0b3f-737f-484a-85b1-811e08771f24 · outbound

This paper cites write newline.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.057222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.057222Z digest=sha256:7eedc2d6741fdba355cfd888606e040c0e2f10578dc59f3e5f39ed4e028c860a

Observation 511f238a-e3aa-4478-8dce-1dd08fed1d9a · outbound

This paper cites Tell me about yourself: LLMs are aware of their learned behaviors.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Tell me about yourself: LLMs are aware of their learned behaviors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.176888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.176888Z digest=sha256:1b7c7d59dad0f60d1a2c8c032307e3d526156b2946b0b54b72762a2767232379

Observation 0b876a4f-14aa-499c-9a3a-99deede786ab · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Emergent misalignment: Narrow finetuning can produce broadly misaligned llms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.246250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.246250Z digest=sha256:b516bfbbc3f5102b4e49388551907c9edb74d9e3ef64035bfc39865e0d15f936

Observation ceccbf25-6bf5-4eaf-b38f-82f0527d3524 · outbound

This paper cites E., Hume, T., Carter, S., Henighan, T., and Olah, C.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability E., Hume, T., Carter, S., Henighan, T., and Olah, C

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.336601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.336601Z digest=sha256:f254eceeda285095bd59577a5ec780c10624ce175c898eb05f18c4c15149c2b4

Observation ce6d09c6-23cc-41fb-a75f-90b362471e25 · outbound

This paper cites Eliciting latent knowledge: How to tell if your eyes deceive you, 2021.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Eliciting latent knowledge: How to tell if your eyes deceive you, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.979962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.419110Z digest=sha256:68054dd89a56e73ee16d0b26d4068293ab5fcb3d71d184273a6d67b068c19506

Observation 8ab3fae9-90a7-4b18-935f-485cfe3fc634 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.533081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.533081Z digest=sha256:8edf01b5ee01f49d269afed747f773871b2dcca850ca21599308cecf4459021a

Observation 5ad1cdce-3136-4de4-8d8a-f201ef897d63 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.657362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.657362Z digest=sha256:455bc9a5873d853fcc61626b56120d118afceb475baadcb24dd54f6506953142

Observation 9020be08-2811-4fb1-9c81-85823fbfd8f9 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Safe RLHF : Safe reinforcement learning from human feedback

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.789023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.758300Z digest=sha256:07c96822cbcc326cc43e281f0c0dec5e64cfd76e81917a0c64ede2e6777d35f8

Observation 2f1fc3d7-4b82-437e-9b8f-da650e9b9df9 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Qlora: Efficient finetuning of quantized llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:05.816404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:05.816404Z digest=sha256:cf87ef87704264f7698d07ca271d6edb1d88fc6c01229bb3eee2e71b450989f0

Observation d8419faf-5ad4-4c82-9089-e10388bd00af · outbound

This paper cites Pal: Program-aided language models.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Pal: Program-aided language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.544027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.888242Z digest=sha256:000b961eb6f1303b4b6b3d7ec56768a52626cb93e163eb174b41d48a327b1d32

Observation f78732ec-7c97-41db-acb6-5b6f3c15ed75 · outbound

This paper cites Gemini 2.5 flash, 2025 a.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemini 2.5 flash, 2025 a

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.346882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:05.969984Z digest=sha256:c1259488eb46a62a6883c267b640f505be346c9e9b7cae97ebad7a150eaf4822

Observation c5f075c8-433c-4927-8594-ea3ab04015d6 · outbound

This paper cites Gemini 2.5 pro preview, 2025 b.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemini 2.5 pro preview, 2025 b

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:10.136279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.061216Z digest=sha256:237e599e4bc58f3b3660aed7d0a95f93c6165cf8ac0c6ed468210935342f3398

Observation b5635376-a01a-460b-8ad0-e949039a0f6c · outbound

This paper cites Alignment faking in large language models.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Alignment faking in large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.156563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.156563Z digest=sha256:9eeebb6038de48efac893592c6fc112ca09f93f44fb7893f11ff561a398a2f46

Observation bc2b0c15-274a-432d-a95d-c421c32a8667 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.249074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.249074Z digest=sha256:005606b9550d93f2856c497812d56029e898d78f6a81f97033d4d3563020a85f

Observation 6dc4628e-ae13-40c4-a7fa-3fa88ab731d5 · outbound

This paper cites and Lee, S.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability and Lee, S

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.952659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.332163Z digest=sha256:2e940f30a838359bdd06ce7463d2209e37f134246d50444fb8ce185d1e1c3d71

Observation 16de2527-ffb6-46a0-801c-457ab90bc719 · outbound

This paper cites u chemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., G \.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability u chemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., G \

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.818736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.413140Z digest=sha256:34da3311953b94890d0f736ee970ad9868b60dfb9d525b1b384c512214c93139

Observation 480279d4-bbad-4752-9cec-e4b7594543c9 · outbound

This paper cites M., Bommarito, M.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability M., Bommarito, M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.674256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.511385Z digest=sha256:d1911dff63b798adeef06f552cdc25cc86889090149ca4d5033db5eaf830b2db

Observation 13023214-b0bf-4312-be60-747afb101da2 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.584693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.584693Z digest=sha256:ba85cb7bd6d27b0ca21472bf35601a106f226bd3b51f70681d7631d3852bc4f6

Observation 4523999c-11ea-4675-8629-f91213f4ebe6 · outbound

This paper cites Auditing language models for hidden objectives.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Auditing language models for hidden objectives

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.682295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.682295Z digest=sha256:91a7ce64b92e298ec423918c9a667aeaff74f7b6b9c718e5843d284a8fc78f19

Observation ac39182b-c4bf-4d4d-9319-78b89bf67f87 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Frontier Models are Capable of In-context Scheming

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.782520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.782520Z digest=sha256:3c6aefafc3b93ad54bbcb0973ff2c57b99266a3e88ca617751ac60759d50491e

Observation 68671958-3532-4304-b1c0-16ed2186ead8 · outbound

This paper cites interpreting gpt: the logit lens.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability interpreting gpt: the logit lens

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.522399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.880431Z digest=sha256:2b320434c80efc8b56e7323fe8d905b0fb548fea2fd614b7233d387638348e0c

Observation efdf8ba3-4011-479b-b770-f283ec91767b · outbound

This paper cites Learning to reason with llms, 2024b.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Learning to reason with llms, 2024b

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.328159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:06.996506Z digest=sha256:b8ee4bd5c8009cc6d0a79facbffb12df42198411561524d931cb56d65ccb2152

Observation 1a02bb6a-e6a2-418c-8618-b60bfa998db6 · outbound

This paper cites Training language models to follow instructions with human feedback.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:07.112022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:07.112022Z digest=sha256:7f3b6d26903299e4047f6517e63a87625fa78fc5472ae35f39cd52f11978bf30

Observation 0f6bfe05-6f84-46d8-91cd-3a1ca2a3bbbe · outbound

This paper cites D., Ermon, S., and Finn, C.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability D., Ermon, S., and Finn, C

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.166740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.229971Z digest=sha256:ab7d9af31c2e128877c741a50051df9d2d3f6b63c47f4ba628abc683b51c2b9d

Observation 49d393c5-a973-41d8-ae91-34bfd73cfeda · outbound

This paper cites LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:40:08.100475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.346882Z digest=sha256:01d6f84581fa9cbb43fe64459f345598cd489d23a151e4f104fe53da1017b355

Observation 7a6f0fd4-9efd-4e5a-82cd-e68e69c925ce · outbound

This paper cites Top 1000 english nouns, 2019.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Top 1000 english nouns, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:09.063228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.459790Z digest=sha256:d2c5c84aa23092c47cc00c23dad16691b74bd9a985e45a4bfe2395fe7cf2e4fc

Observation 87906699-feee-48b8-9fe9-90d16caa5973 · outbound

This paper cites Large language models can strategically deceive their users when put under pressure.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Large language models can strategically deceive their users when put under pressure

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:40:08.902763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.539751Z digest=sha256:29b69a6b4b74279911f9496bb490e41b801f4c0f91670bf188346427e3e66a9a

Observation 12d7963d-ac67-4556-8d4c-d04a4424e592 · outbound

This paper cites an unresolved cited work.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:07.649664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:07.649664Z digest=sha256:58dba48e502485f9f38f1c9c820e3a1e4661d70e11ddf9dcf5a5932b3997da99

Observation 1e8a314b-57e8-4e7d-923c-ea400b0faca4 · outbound

This paper cites an unresolved cited work.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:40:08.587883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:40:07.773307Z digest=sha256:7a0bb3f83e997857d30ae568ffdebfdc12c8d56fef44c2b4f307765ba748f8c9

Observation 6ee307d4-e2d9-4355-ad8b-597a447b6ec0 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability N., Kaiser, ., and Polosukhin, I

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:07.861843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:07.861843Z digest=sha256:79cee4bf708f50daafcb89691fb3a216463ae65860fa828e775703e1144fb272

Pith citing papers

Observation b4bd3ba3-9c28-48ce-82c3-e71a7d58af84 · inbound

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy cites this paper.

Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:20:51.989267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:20:51.989267Z digest=sha256:d611de8be7fffbb4ebfe55d26f3a6b3c781b5fd71a45d539acd0d7a1aade6192

Observation d7e6064b-8eca-414d-913f-f9b0a427f6ea · inbound

DECOR: Auditing LLM Deception via Information Manipulation Theory cites this paper.

DECOR: Auditing LLM Deception via Information Manipulation Theory Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:28:05.444725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T06:27:10.445757Z digest=sha256:3e245a9bedf6cd59d2f02a331f35fd3643bcfdeff2734193f219421b4a8e6e23

Observation 71882747-c6a5-4999-bdb2-b6a74edf1618 · inbound

Building Better Activation Oracles cites this paper.

Building Better Activation Oracles Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:24:45.078140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T14:19:37.262201Z digest=sha256:fcf5faa91399bfd707ecf8f85f3eeddd8818a04ec79e47b29d7bded024400851

Observation bf2dbefc-e764-4a38-9b8d-6660d1c8a799 · inbound

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms cites this paper.

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.307763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:51:16.969884Z digest=sha256:9beb02ead55c39a61a64889e6913b1865a55bc8a097e445fcaacb457d4903c64

Observation a44a5147-ccd4-4129-a339-7880f3ca094a · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.997805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:e678be8cffc2b1f0233478de7576ade12d11826197294fdde5c431808db1fb72

Observation 3a8b2a06-b7ed-42ab-b5fc-883db645c7ee · inbound

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo cites this paper.

"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:58.191057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T13:11:53.361408Z digest=sha256:883dc5da9756d33eb82445cc2a9da1cae4ffeb1aae8fe19ed7aa4a07bc09863f

Observation b9cd5712-4558-422c-9c58-77fc742c2630 · inbound

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology cites this paper.

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:07:08.173995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-02T15:57:48.589980Z digest=sha256:b77e685ee3f8cc436b9ffe9db6949d2ba3853c6178ca22f8d6adc765d352cc17

Observation 76e1f9be-c3ec-4f0e-9a94-c8040b4a08fe · inbound

MUX: Continuous Reasoning via Multiplexed Tokens cites this paper.

MUX: Continuous Reasoning via Multiplexed Tokens Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:03.078595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:43:03.078595Z digest=sha256:4a1c101e9bd47de31e7b1dfdba57b66ef106ee390d9fca4f640a047b0a45b8af

Observation df4e3907-474c-4909-9ba9-029b28afc797 · inbound

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles cites this paper.

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-30T23:56:50.935284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T23:56:50.935284Z digest=sha256:9ede6d0979f419d5d493cd00226a860d6ee167081773957499355076932defa5