Pith. sign in

Paper Citation Record · LEDGER

Representation Engineering: A Top-Down Approach to AI Transparency

As of 31 July 2026, this Paper Citation Record lists 23 of 23 outbound references and 100 inbound Pith citation observations for arXiv:2310.01405.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.01405 v4

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 123 of 123 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 100 of 336 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T14:15:29.705679Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T15:37:20.411640Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34e1c715-38b8-4c2a-b10c-199cb9b6b92d · outbound

This paper cites Language Models are Few-Shot Learners.

Representation Engineering: A Top-Down Approach to AI Transparency Language Models are Few-Shot Learners

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T18:11:40.992027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:ad28935510b731a091db4e84924a58af8cebb261fc45c180ba1e402587bbe3fb

Observation c4953bea-638d-4984-a414-c8c11b18ccc2 · outbound

This paper cites doi: 10.18653/v1/D17-1082.

Representation Engineering: A Top-Down Approach to AI Transparency doi: 10.18653/v1/D17-1082

Reference 2

Resolution
verified exact
doi, observed 2026-05-10T18:11:40.995924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:ddc41715a7fa07543fec9636f17200ca3f725cf1c9058e91fcb7cd07bb7119d6

Observation 90ee26c9-7169-4f2f-b67f-481820211c1d · outbound

This paper cites doi: 10.18653/v1/n19-1421.

Representation Engineering: A Top-Down Approach to AI Transparency doi: 10.18653/v1/n19-1421

Reference 3

Resolution
verified exact
doi, observed 2026-05-10T18:11:40.999844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:8a28114cd762d6fdea4ef1efbad3c69d99a8ea85701cc78c7c061a0ca65e0867

Observation 897a3a8d-fcf7-490d-8585-7d424ea65e5a · outbound

This paper cites Love” and “Hate.

Representation Engineering: A Top-Down Approach to AI Transparency Love” and “Hate

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.003358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:8e8de53c991e83274e6d45c0d5d2b03dcaa0b4ed9056aac48588a0472be5dc21

Observation c543172b-36f0-472f-87a4-6b3d24a69235 · outbound

This paper cites We take the top PCA direction that explains the maximum variance in the data X D l.

Representation Engineering: A Top-Down Approach to AI Transparency We take the top PCA direction that explains the maximum variance in the data X D l

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.006338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:3156f1e509166aa5b0264f502ca774b75bb60ca4f1b13cc07d339a2f36778d02

Observation cbe1f478-c849-4874-b4c2-04b03c9c59a7 · outbound

This paper cites We take the difference between the centroids of the two clusters as the concept direction.

Representation Engineering: A Top-Down Approach to AI Transparency We take the difference between the centroids of the two clusters as the concept direction

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.009220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:a518f970413d5c6810369549ac39e2d808b14aca1fdb4c2352309950d3bb8b63

Observation 6e7a9a9a-c9ce-460c-bb82-a1c57dea0247 · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.011815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:1f765e1055dc1e721722766a24949abb40807e2eeb3491c4fb961d82d478d208

Observation f10fe97c-1265-4fc8-90ad-913e52d61bec · outbound

This paper cites I made a mistake and copied my friend’s homework. I understand that it’s wrong and I take full responsibility for my actions.

Representation Engineering: A Top-Down Approach to AI Transparency I made a mistake and copied my friend’s homework. I understand that it’s wrong and I take full responsibility for my actions

Reference 8

Resolution
malformed identifier
raw_fallback, observed 2026-05-10T18:11:41.014706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:7bc4372abc406f2eb2482de0a1299a6a84066347bdbbe3e51f311cc4f477031b

Observation edfae63c-b5fa-46b1-94dd-f80069a50e8f · outbound

This paper cites psychopathic.

Representation Engineering: A Top-Down Approach to AI Transparency psychopathic

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.017500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:f5914c6ce2e8057efc804c5d271d2ee3f68ea63d1df5b2fa71ed0b28cc0358a2

Observation a333616f-b3f7-46fa-87c6-3c5ada89566d · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.019986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:871aa3250988ee6f0f849072fa028730c3705a7af089b33673ff7d87a3011e2a

Observation 01c27a0c-0562-439c-a4e6-57e5f76f7571 · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.022808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:e3b3d0b1b9fce782dcd0d24ab4c48515659f47c4001a022b6601b6c1b83a0269

Observation 68d291ee-1aec-4b21-a3dd-6f0518f16ed5 · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.025401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:814bfeef940d15a28fba7fb4ee79a242ba36551295314d4835458f76ef5b6ecc

Observation f418cdc6-15c3-461d-b8b8-02c15a24a597 · outbound

This paper cites Do the findings rest on strong theoretical assumptions; are they not demonstrated using leading-edge tasks or models; or are the findings highly sensitive to hyperparameters? □.

Representation Engineering: A Top-Down Approach to AI Transparency Do the findings rest on strong theoretical assumptions; are they not demonstrated using leading-edge tasks or models; or are the findings highly sensitive to hyperparameters? □

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.028176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:15a15d41bba50e239ef29982ba3bdb4a048c2a2dd5de7862d73a60d5604e76f6

Observation 8a662b95-5a9d-43fd-8961-d73b25588565 · outbound

This paper cites Is it implausible that any practical system could ever markedly outper- form humans at this task? □.

Representation Engineering: A Top-Down Approach to AI Transparency Is it implausible that any practical system could ever markedly outper- form humans at this task? □

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.030691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:a56db73428fced26a20df009b72a6d71c19b9a67577126ff0428090ef745272d

Observation 9f60eb8d-c717-45a4-bde5-47e66297f023 · outbound

This paper cites Does this approach strongly depend on handcrafted features, expert supervision, or human reliability? □.

Representation Engineering: A Top-Down Approach to AI Transparency Does this approach strongly depend on handcrafted features, expert supervision, or human reliability? □

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.033204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:384a10be6de2b382f8e1e6a92d3e2d8fa4215a0527cc4206a47020374d148020

Observation c6a435fd-2828-44fa-a69f-6442515ab9f2 · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.035435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:5b4059cefb6a9d1df6c67aa74b6bb2f7ac14b7626121620e00f1ae0ee31b7714

Observation 24b981ee-bce1-4524-844b-965e08d46b4e · outbound

This paper cites How does this improve safety more than it improves general capabilities? Answer: This work mainly improves transparency and control.

Representation Engineering: A Top-Down Approach to AI Transparency How does this improve safety more than it improves general capabilities? Answer: This work mainly improves transparency and control

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.037925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:b74c05fb356d9bd3bd15ed8b6776733d53169872b6498bf1d93de786e0de286a

Observation 84ff59e6-00d6-4ca0-93bc-eed2f3dbe05f · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.040111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:a1c1ed4b9cd129b6e774f7e887014e2f5453554c64a02c74ae6df41740466b3d

Observation 273778c6-aadf-4923-8a35-46119b923b53 · outbound

This paper cites Does this work advance progress on tasks that have been previously considered the subject of usual capabilities research? □.

Representation Engineering: A Top-Down Approach to AI Transparency Does this work advance progress on tasks that have been previously considered the subject of usual capabilities research? □

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.042466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:7f07f8da3699260dad0f3b4788ac2e72b13deb70cd8732f096a1f2043d392d7c

Observation 84f757df-78ba-4bbd-9e82-206b2dd2d0dc · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.044790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:3bb8ea69debd19cc2081717fda2f91e3e2109da96dccf395523a17498cfd46bb

Observation 74e244a9-be65-4193-b51d-deb26d4fc3c3 · outbound

This paper cites an unresolved cited work.

Representation Engineering: A Top-Down Approach to AI Transparency Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-10T18:11:41.046954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:873d60858fb27fec8d93a9bde46aa0a7dd203543c3d858e98c4cdae1a2a7248b

Observation b091299d-93f6-4c2f-b2c7-dff173ec1c55 · outbound

This paper cites Does this advance safety along with, or as a consequence of, advancing other capabilities or the study of AI? □ E.3 E LABORATIONS AND OTHER CONSIDERATIONS.

Representation Engineering: A Top-Down Approach to AI Transparency Does this advance safety along with, or as a consequence of, advancing other capabilities or the study of AI? □ E.3 E LABORATIONS AND OTHER CONSIDERATIONS

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.049442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:9144606f983cacec06462c069f0e3efd9ef1f72f56f8a7ca5e6edbe0398b4276

Observation 61f2fb19-62ef-440f-90e4-2231b7923c0f · outbound

This paper cites complex and fragile.

Representation Engineering: A Top-Down Approach to AI Transparency complex and fragile

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:11:41.051802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:11:40.970904Z digest=sha256:c07d6464989926f90c5b1cc4b993c56d391851a7b842c804a3520ec36300f658

Pith citing papers

Observation a8d08171-95b0-4a21-aff0-fa279eaa9dcf · inbound

AI Alignment: A Comprehensive Survey cites this paper.

AI Alignment: A Comprehensive Survey Representation Engineering: A Top-Down Approach to AI Transparency

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:28:49.016239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-17T14:28:48.987140Z digest=sha256:1eaa71fcf8e0f3f25401e23c09ec6e5a6104733244293381a8781a65d4ef1074

Observation c909f869-1d84-4ec2-aab2-17f6a6155805 · inbound

The Linear Representation Hypothesis and the Geometry of Large Language Models cites this paper.

The Linear Representation Hypothesis and the Geometry of Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:43:30.838860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-11T21:43:28.274354Z digest=sha256:af1bc6a19f586861054dc647bb86d4da9ffc1996f86c5145251f833417770c49

Observation 7739f3cb-6554-467d-b692-df57b12dbaeb · inbound

Steering Llama 2 via Contrastive Activation Addition cites this paper.

Steering Llama 2 via Contrastive Activation Addition Representation Engineering: A Top-Down Approach to AI Transparency

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:37:21.300128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-11T20:37:20.408376Z digest=sha256:b20ea15607e3251108c90eb6bbf0f60b69fac0ac56e1c578ebfc33f64e00371d

Observation 72bf91dd-6833-45d9-ab54-708e9c017267 · inbound

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models cites this paper.

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:15:10.727523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-13T13:15:10.632115Z digest=sha256:338740a9bbaafe4f91f1f044be0e4099c6fafd16de11342dbbb58ec6e2246714

Observation 762be8d1-2052-45c4-8f55-2c7d8c9da9fe · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Representation Engineering: A Top-Down Approach to AI Transparency

Reference 207

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:47:56.176618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:8997134b430c29976340b218f568ef272d6ee923128305da1b9de7daf1359258

Observation 974adf50-0bce-4cc9-a758-21066380b6a0 · inbound

Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs cites this paper.

Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs Representation Engineering: A Top-Down Approach to AI Transparency

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:52:02.604208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T00:52:02.421389Z digest=sha256:1ce7281a6caeb28c0467ff39727d915dae1e124c232109bbdbce2d7eafbbcf04

Observation a003612c-133c-4085-907e-c3e67ca4a966 · inbound

Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct cites this paper.

Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct Representation Engineering: A Top-Down Approach to AI Transparency

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T19:53:23.230920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-23T19:51:14.630756Z digest=sha256:d67794b060bdf5073f47baab07a854a1ce9b1d731cb54296e090d4644ed2c770

Observation f996f9eb-11eb-4278-843d-81c8ac638337 · inbound

Open Problems in Mechanistic Interpretability cites this paper.

Open Problems in Mechanistic Interpretability Representation Engineering: A Top-Down Approach to AI Transparency

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:30:06.022244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:3b798d2ee09f9f78fd1813cf589c04a2152d5d74bc9cbbc027fa79c982b23375

Observation 47bc9731-5628-4577-ab9e-28c298906cd5 · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Representation Engineering: A Top-Down Approach to AI Transparency

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:27:21.335303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:68437b207d234f8443036353dd876965f2e1b3e08d37ea7a9228a5857a675689

Observation ca3411fb-7f32-44d6-a529-5f2db1051007 · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:24:13.007302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:14ab1de91181d481339c6d81664a6e85995ccb2e94398d597fdfdcc6f1eb4002

Observation 467663c3-9f95-4987-b99c-0198592831fd · inbound

Secure LLM Fine-Tuning via Safety-Aware Probing cites this paper.

Secure LLM Fine-Tuning via Safety-Aware Probing Representation Engineering: A Top-Down Approach to AI Transparency

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:11:35.874410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-22T13:07:09.402763Z digest=sha256:9bba419ed179dea924d7c0efd749f238895117888cfa2fba5f9ea87acc456d77

Observation cc9d089b-3a7b-41d9-9773-441baa675efe · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment Representation Engineering: A Top-Down Approach to AI Transparency

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:53:03.601789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:4d72673e44974288866b97291d1934d2fff3794f2cb7b2da01ef23749d536fec

Observation c0284196-7fd1-4e22-8b2e-f939c54795d0 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Representation Engineering: A Top-Down Approach to AI Transparency

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:37:15.901459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:0dd164bf9001a1cb466e8efed42ded1ac7662d726bd1528777fc53a6dc85249f

Observation 93e4d6bd-b054-4117-b748-264fab7490fc · inbound

SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention cites this paper.

SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention Representation Engineering: A Top-Down Approach to AI Transparency

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T09:37:14.191045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-19T09:34:24.194855Z digest=sha256:bf7614d37ae65bc5580c3efabdec53ea29ec5a2ab50f316c8603c3b45aedcf15

Observation 767a7f2d-419b-4a88-b238-07a6b491e0db · inbound

AI Feedback Enhances Community-Based Content Moderation through Engagement with Counterarguments cites this paper.

AI Feedback Enhances Community-Based Content Moderation through Engagement with Counterarguments Representation Engineering: A Top-Down Approach to AI Transparency

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:27:05.576965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-19T05:23:00.858183Z digest=sha256:f26078ba282151b724c086f220c75ad394c4cdc7b96842480b1c8a1d18fa7431

Observation faa59524-a754-4429-91de-85cb5b050602 · inbound

Similarity Field Theory: A Mathematical Framework for Intelligence cites this paper.

Similarity Field Theory: A Mathematical Framework for Intelligence Representation Engineering: A Top-Down Approach to AI Transparency

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:41:30.673119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T14:36:49.543497Z digest=sha256:6431037d04fac27d67638375b05c93ee4bd0d02d1ca68baac0394dfc46695fa3

Observation 9d68e2ae-6f90-4bc3-b11f-fe9cec573840 · inbound

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models cites this paper.

Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 26

Resolution
malformed identifier
local_arxiv, observed 2026-05-21T21:45:40.623814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-21T21:44:36.351517Z digest=sha256:75c2692ff00bf2e8958eb0e40040f7390ae6c289715ec1048fe0f02a3ded8483

Observation e86d5acf-2b40-4fbb-b48f-c9f716086568 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs Representation Engineering: A Top-Down Approach to AI Transparency

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:55:38.026427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:628b616ce3b24abe4468edf49fb340ae787e3626ce15e260698b0a3718f5c708

Observation 9c81199a-9595-447f-98ff-0752ae4c49d0 · inbound

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations cites this paper.

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations Representation Engineering: A Top-Down Approach to AI Transparency

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T18:50:30.288269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-21T18:46:04.926179Z digest=sha256:b35fb3a351d6a70cbd084cafb9e23d244b769ede8d1af25f31eeef6ccca72694

Observation d506ccca-0558-4d99-a37f-5f0ee8994d3b · inbound

The Impact of Off-Policy Training Data on Probe Generalisation cites this paper.

The Impact of Off-Policy Training Data on Probe Generalisation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:30:11.744338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:a7fdf938e8b5bc4d945f3d308ac30ada24b9c41ce864f58bed098334d43d9dc4

Observation 67e7d001-9de5-48e8-b2af-5f9bedfcfbdf · inbound

Sparse Concept Anchoring for Interpretable and Controllable Neural Representations cites this paper.

Sparse Concept Anchoring for Interpretable and Controllable Neural Representations Representation Engineering: A Top-Down Approach to AI Transparency

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T22:18:36.673480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-16T22:14:40.293207Z digest=sha256:82be56e7632a5b4d842a11f1e3c2a902e07bc71cf69ad63e6b43f5f38fe02db4

Observation 7ad837fd-5640-4b52-9a39-1805d0bafb83 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Representation Engineering: A Top-Down Approach to AI Transparency

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:17:36.626928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:cb626012550d33ffabe80f1d65d757d2c1b032951c6f98fe9752973c11393611

Observation 40248c51-1738-4687-8d8b-9f4f026f8451 · inbound

Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models cites this paper.

Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:07:20.474413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-16T05:06:24.973439Z digest=sha256:d4bbc843181afceb3e1f36650b46b1237414c259a75c9528ed11a0cf9fc5fef9

Observation b558ed84-735a-4a84-8cec-4ed510f8d4c7 · inbound

Activation Steering for Accent Adaptation in Large Audio Language Models cites this paper.

Activation Steering for Accent Adaptation in Large Audio Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-15T14:15:29.705679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:15:29.705679Z digest=sha256:8ba59d3eb6f85f2045b0e03fff9349e2296122026f07d74172d1c9df968018fa

Observation 6233aaca-4963-43cd-bdc8-e40c11749af3 · inbound

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges cites this paper.

CreativeBench: Benchmarking and Enhancing Machine Creativity via Self-Evolving Challenges Representation Engineering: A Top-Down Approach to AI Transparency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T22:34:21.790260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:34:21.790260Z digest=sha256:526b3988e34cff76a10fcb928e29969300da64a88e40ffce392a02b67b6ca99c

Observation a7e97ba6-73e4-4bb9-8dbd-972fe09300d2 · inbound

Why That Robot? A Qualitative Analysis of Justification Strategies for Robot Color Selection Across Occupational Contexts cites this paper.

Why That Robot? A Qualitative Analysis of Justification Strategies for Robot Color Selection Across Occupational Contexts Representation Engineering: A Top-Down Approach to AI Transparency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T16:04:03.659837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:04:03.659837Z digest=sha256:f21cab91f7d360cd9b4938c1fd56155a34c7309eafae444a912358d84ead2843

Observation 16a9e383-eb52-4b32-8cb9-5b01e20e00da · inbound

Critical Damping as a Momentum Schedule: Multi-Seed Validation, a Hybrid Recipe, and an Exhaustive Negative Result on Surgical Layer Selection cites this paper.

Critical Damping as a Momentum Schedule: Multi-Seed Validation, a Hybrid Recipe, and an Exhaustive Negative Result on Surgical Layer Selection Representation Engineering: A Top-Down Approach to AI Transparency

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:38:00.031448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-14T21:37:07.118251Z digest=sha256:0b4670ff4827a521134a43f7a0e2469bcba0ac56894f128b71f2983899331b0f

Observation 7e85a505-08d8-4b05-8b5b-2bbe3c106fd7 · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Representation Engineering: A Top-Down Approach to AI Transparency

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:fc6b99cf496c7b8f45f639fcf7a5f821de286a5004c7ac234f87100d76935267

Observation 90631e93-3173-49a6-a13d-66d31491d0cc · inbound

Dual Implications of Quark Mass Hierarchies to Flavor Structure cites this paper.

Dual Implications of Quark Mass Hierarchies to Flavor Structure Representation Engineering: A Top-Down Approach to AI Transparency

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T13:46:54.451163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:46:54.451163Z digest=sha256:94df523af0518d3b8ee021921089ed3543f0cfc924b0a3c0d7b7237ca24da5a2

Observation d12cec21-60f2-4a68-9b63-79135c52db13 · inbound

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens cites this paper.

Steerable but Not Decodable: Function Vectors Operate Beyond the Logit Lens Representation Engineering: A Top-Down Approach to AI Transparency

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:03:20.239459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-13T20:59:02.824707Z digest=sha256:4ba604ab7681e9b4f4600d7ea28d3cab6f090c284ed71943c0deb6c614e7f970

Observation 68592d14-f536-4c61-b56c-578299073119 · inbound

Enhancing Multi-Robot Exploration Using Probabilistic Frontier Prioritization with Dirichlet Process Gaussian Mixtures cites this paper.

Enhancing Multi-Robot Exploration Using Probabilistic Frontier Prioritization with Dirichlet Process Gaussian Mixtures Representation Engineering: A Top-Down Approach to AI Transparency

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T13:35:02.428552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:35:02.428552Z digest=sha256:0fe266ad4fc4f327edf2565ae163a6a4d73be486a4ba38c9cdaa5f503edc71ac

Observation 7483e0d1-8b6b-4e8c-a63c-7fd4d199cffe · inbound

STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models cites this paper.

STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:18:13.365221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-13T20:16:04.999705Z digest=sha256:6f4d5c88afb9b6555db8df1e69fe2dd22a2ac553f0e2da1a595d4d5e204f77ad

Observation f5d6f474-d6f1-4f5d-a118-bf05a504d1b6 · inbound

Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control cites this paper.

Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control Representation Engineering: A Top-Down Approach to AI Transparency

Reference 4

Resolution
malformed identifier
local_arxiv, observed 2026-05-13T20:18:13.410647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-13T20:15:41.756907Z digest=sha256:1270257ae344ecb4062178660ea81aae7426b279060e5caf25bb2703dc34ee45

Observation ee097d1e-6886-4343-9699-5ad26e50246a · inbound

The Democratic Ontology Deficit: How AI Systems Fail to Represent What Democracy Requires cites this paper.

The Democratic Ontology Deficit: How AI Systems Fail to Represent What Democracy Requires Representation Engineering: A Top-Down Approach to AI Transparency

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:53:00.117396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-13T16:50:53.440923Z digest=sha256:0540226829ecb4e63a374098030bfaf0ea0bff722bf51aa22789599c111462c5

Observation 9b8d3568-179c-4390-98f6-f58858ce0fd5 · inbound

Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment cites this paper.

Where to Steer: Input-Dependent Layer Selection for Steering Improves LLM Alignment Representation Engineering: A Top-Down Approach to AI Transparency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T12:15:38.881394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:15:38.881394Z digest=sha256:a1eb7d9ec5c239c9084c1f14161a305f53b90909c160bd0641540a02cc8b55ab

Observation db8ad889-984a-4ddc-850b-8fa31ba755a9 · inbound

How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models cites this paper.

How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:10:50.518999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T20:09:51.378878Z digest=sha256:efd6400d59180d19f67b977c1cdbc6d214d27954aac15cb43f661c0b5063b9e3

Observation 2d840371-9a51-4e69-8735-c9201f92d8aa · inbound

Sparse Autoencoders as a Steering Basis for Phase Synchronization in Graph-Based CFD Surrogates cites this paper.

Sparse Autoencoders as a Steering Basis for Phase Synchronization in Graph-Based CFD Surrogates Representation Engineering: A Top-Down Approach to AI Transparency

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:19:31.717017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-14T22:18:22.476746Z digest=sha256:32d0de01d187ff8e7c422cde5be1c16492608f391477ba99181e411ac35acf16

Observation c40bc05c-648e-4552-9c3d-1ea7712e242e · inbound

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment cites this paper.

The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment Representation Engineering: A Top-Down Approach to AI Transparency

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:15:56.188766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:36:44.401045Z digest=sha256:d2b1802724b40e0cde828968a92892ffe512180d6749409821250674740892de

Observation 56bf7d92-05bc-4ca5-bd16-d69016395486 · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:25:52.313537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:31:17.182481Z digest=sha256:fa047ba1774dcbbbf19edf5deed9984b9335c09c76a6c0f930970fc2890f89a2

Observation 62121618-702f-4216-baff-ec075ff1d25c · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:25.123020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-12T04:28:05.507808Z digest=sha256:6f9caaf91b576cd1db4d0df93fd810f99e8c148ebf390111403fc180765065e2

Observation 0cf4070d-27d7-4e8f-a76b-b6a47ca4b29e · inbound

Emotion Concepts and their Function in a Large Language Model cites this paper.

Emotion Concepts and their Function in a Large Language Model Representation Engineering: A Top-Down Approach to AI Transparency

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:35:58.213915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:03:52.210931Z digest=sha256:f67011f8f976483a06cd28c8ded4112609c7b24098d1ffc5dc102558938a6c38

Observation 23247e79-193c-49f3-8b41-a7cc10254a8f · inbound

Beyond Social Pressure: Benchmarking Epistemic Attack in Large Language Models cites this paper.

Beyond Social Pressure: Benchmarking Epistemic Attack in Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 24

Resolution
malformed identifier
local_arxiv, observed 2026-05-11T00:30:55.137260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:27:43.273410Z digest=sha256:9698cc424272c36b124beaf1a413c4c85963d7c633b0ffc536daffbd03222ae0

Observation 060b2fb4-f881-4a2e-99b3-b591ed7607e0 · inbound

Linear Representations of Hierarchical Concepts in Language Models cites this paper.

Linear Representations of Hierarchical Concepts in Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:45:57.607911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T17:00:04.457449Z digest=sha256:b892bb1960fbf700a29652275d2445d06d3a7d5941a8eeabcafb1a1e344fc663

Observation e7017615-6ff0-46cf-9956-da1f2311ae60 · inbound

Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models cites this paper.

Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:55:57.478701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T17:53:35.970156Z digest=sha256:29881ddbc9d44a384f965745d00e75bbcf05e06fc1e840a7e96bb8c59337cd78

Observation 7420bd99-1623-4271-b4c9-9dc983399cd6 · inbound

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal cites this paper.

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal Representation Engineering: A Top-Down Approach to AI Transparency

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:05:59.837702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T17:47:43.189696Z digest=sha256:9ed7f412626cc46511f47be1c5a3275627bf505b2622584e0e978037630db26e

Observation a79a4202-5578-4dad-b4f2-29b5bde62af8 · inbound

Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest cites this paper.

Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest Representation Engineering: A Top-Down Approach to AI Transparency

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:26:02.224964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T17:11:13.189660Z digest=sha256:181a419dd1303340b6b29e8c6ad3070d9be5ff570cffbd53a8b9ade800af7257

Observation 51ea0f1b-925a-4579-a7d4-90f811fad919 · inbound

Spectral Geometry of LoRA Adapters Encodes Training Objective and Predicts Harmful Compliance cites this paper.

Spectral Geometry of LoRA Adapters Encodes Training Objective and Predicts Harmful Compliance Representation Engineering: A Top-Down Approach to AI Transparency

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:21:01.623782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:10:42.847854Z digest=sha256:039de868fde142dec821e357ece2b6e1cd9cf272fb95b75c2b4bc7022ddb4bd1

Observation 4a9b941b-72e5-4e68-858c-319159b3be48 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Representation Engineering: A Top-Down Approach to AI Transparency

Reference 134

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:35:57.476712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:0e92248e62a7c3398def4a02b22e74f18083340a0c966dc3d6c854ca4b11b525

Observation d7b4aecb-48d8-4a4b-9fc7-57646ee8dd16 · inbound

SHIFT: Steering Hidden Intermediates in Flow Transformers cites this paper.

SHIFT: Steering Hidden Intermediates in Flow Transformers Representation Engineering: A Top-Down Approach to AI Transparency

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:20:55.675839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:33:59.134475Z digest=sha256:dc25ee97eebdacfc1e92a31aab4b2a105530ac68fa1b72c824a020be0314e6dc

Observation dd40cf67-790e-4663-83db-6f2c8a1485c4 · inbound

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems cites this paper.

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems Representation Engineering: A Top-Down Approach to AI Transparency

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.947647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T16:07:31.602378Z digest=sha256:181bac81d2b10f18697a94fb0cd048d1a4a5b68ae4429e38049b2e6f83c755fd

Observation 8ac08129-784d-4f9b-abae-ae8e2a641f50 · inbound

Disposition Distillation at Small Scale: A Three-Arc Negative Result cites this paper.

Disposition Distillation at Small Scale: A Three-Arc Negative Result Representation Engineering: A Top-Down Approach to AI Transparency

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:00.987401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T16:09:21.845152Z digest=sha256:0cb791e53f1782f1ac88dd59ced41247acce029b849284e4f039f95975cb40d0

Observation 7936659b-cb27-44e9-80a0-b1a03ceb3dd6 · inbound

ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems cites this paper.

ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems Representation Engineering: A Top-Down Approach to AI Transparency

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:50:59.101522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T16:27:18.594869Z digest=sha256:7a8b3ea1230b240ed672e72814133ae98e2450add02e255aae084b8246919e16

Observation bc9c2adb-b9c9-462e-ab72-d5f182d1dbe0 · inbound

Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering cites this paper.

Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering Representation Engineering: A Top-Down Approach to AI Transparency

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T09:36:07.724906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T15:55:12.611847Z digest=sha256:ee7a9f12108760da607bbf6338483072ffee050714d23b31ec58f730052d2c8a

Observation 05cf5f90-b603-4160-b5a9-809c7e5cf764 · inbound

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior cites this paper.

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior Representation Engineering: A Top-Down Approach to AI Transparency

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:17:58.775434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-14T21:16:40.430384Z digest=sha256:bd69bd8766d681b558af826a21558adfecc69425c34be6f613bdaeecae8be55b

Observation e4737c71-35e5-420a-b5dc-79642019ce16 · inbound

A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models cites this paper.

A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 54

Resolution
malformed identifier
local_arxiv, observed 2026-05-11T08:55:59.826347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T16:26:32.447186Z digest=sha256:361bdba85f1c2cf8ea4150c4cf0788cb0123a1419fd7aec878d9002e5ea8e0ac

Observation 6eedbe06-41de-41b0-a9b8-0e54251012b2 · inbound

A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models cites this paper.

A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 59

Resolution
malformed identifier
no resolver link, observed 2026-07-12T20:55:34.837390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:55:34.837390Z digest=sha256:2cb7e428ac101065b95b8d35874f9759106655755e310f94fe1b3a14eb5191e9

Observation 2ebbd697-b150-4a73-9a23-46b4d6672eb9 · inbound

The Cognitive Circuit Breaker: A Systems Engineering Framework for Intrinsic AI Reliability cites this paper.

The Cognitive Circuit Breaker: A Systems Engineering Framework for Intrinsic AI Reliability Representation Engineering: A Top-Down Approach to AI Transparency

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T13:56:05.550669Z digest=sha256:16672fa0b3c4578a24d2df050d5963e09acda334e0eb0c3c3e20f1d8ec7eea51

Observation de151e43-9f2c-4ffd-b4e8-923961875a60 · inbound

Weight Patching: Toward Source-Level Mechanistic Localization in LLMs cites this paper.

Weight Patching: Toward Source-Level Mechanistic Localization in LLMs Representation Engineering: A Top-Down Approach to AI Transparency

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T13:31:36.370422Z digest=sha256:e475fff69be37f20d045aef6ef7d5eec614580e1a5203bc42e3a1f035c6c685d

Observation 43e10c68-cf43-486b-bf15-28e57de235ce · inbound

Geometric Routing Enables Causal Expert Control in Mixture of Experts cites this paper.

Geometric Routing Enables Causal Expert Control in Mixture of Experts Representation Engineering: A Top-Down Approach to AI Transparency

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T12:53:48.734715Z digest=sha256:150e2f115dc205a14bbfc00fff72bdcc67e67aca1077bd50fde9cda2ab29f26a

Observation 9a3a1d76-6b5f-4ac4-9192-5dd6360d0e91 · inbound

Psychological Steering of Large Language Models cites this paper.

Psychological Steering of Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T12:44:46.997345Z digest=sha256:879f2514d29c75df7b866d1e9a4755f38418e821c0109cd1cef6618c8b541ceb

Observation 6a4b7f0e-eb79-44a2-9a73-12f5ea1a64f3 · inbound

Mechanistic Decoding of Cognitive Constructs in Large Language Models cites this paper.

Mechanistic Decoding of Cognitive Constructs in Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T12:13:25.350059Z digest=sha256:33302725e5d0828fee3edc0f18baaf4ade967bb2ee9686d6922faa0833d6018e

Observation 5d8926e8-ca78-4354-8e95-5ed73134ba16 · inbound

Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation cites this paper.

Hallucination as Trajectory Commitment: Causal Evidence for Asymmetric Attractor Dynamics in Transformer Generation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T11:02:21.049008Z digest=sha256:bdea1ba999963d9a5db267b9fa68ca29863bebaa8f1207d614dc32a5e11bc229

Observation 66628089-06b3-416f-8bcb-96aa793c02a5 · inbound

Predicting Where Steering Vectors Succeed cites this paper.

Predicting Where Steering Vectors Succeed Representation Engineering: A Top-Down Approach to AI Transparency

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T10:59:30.755424Z digest=sha256:1ad03f5c30e05cfc0844516a2b4053147860755955af2e7c6d70def3070097a3

Observation 9c082272-1be1-4588-b6f7-dd7a60376495 · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Representation Engineering: A Top-Down Approach to AI Transparency

Reference 220

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:5f850ed3dd7c644cc613611f60b262a95b86af46f9c3095ac84ebf45f313eeca

Observation cc261d1e-a70a-42bd-9c21-69c7a0e6cf6b · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Representation Engineering: A Top-Down Approach to AI Transparency

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:be76e5042434b65ed8c633fc5680dc80da1ca28a70c76fe55e4bec88fb150c51

Observation 39231c5d-74c8-4e44-91c0-f54a8409ab28 · inbound

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models cites this paper.

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T05:23:15.746795Z digest=sha256:f1ad43ef20f63748e968462accb7999ba77827ef7571c6348f1b009c7845cebd

Observation a6f3544b-6c7d-4296-803b-c6dbc7172063 · inbound

State Transfer Reveals Reuse in Controlled Routing cites this paper.

State Transfer Reveals Reuse in Controlled Routing Representation Engineering: A Top-Down Approach to AI Transparency

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:56:29.798826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T04:27:01.220899Z digest=sha256:7dbcd3edcd875c7a331a51f9394804f747d3fa3267e4724be8ed1147d20f9154

Observation 8ab12e75-1eb4-4ca3-942a-7fb5d2f8b5d8 · inbound

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams cites this paper.

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams Representation Engineering: A Top-Down Approach to AI Transparency

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T04:34:24.866119Z digest=sha256:35f71f1132a27011016e210a3547484f784e7331f4b1bc4c6b31f1e5bc6b3ee5

Observation bff81ab4-1b5a-4484-93fe-370205cf9b89 · inbound

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams cites this paper.

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams Representation Engineering: A Top-Down Approach to AI Transparency

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T00:51:15.525177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-12T00:49:53.613850Z digest=sha256:9ec305a76251a5b536599b927ece8270a807be1ebf4af56f83ad87d6aadb9be2

Observation d71f28f5-957f-4289-83f1-f95a078c91f9 · inbound

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control cites this paper.

Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal Control Representation Engineering: A Top-Down Approach to AI Transparency

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:01:04.024089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T02:31:07.932802Z digest=sha256:69442da7df6d90dbd7ca52144c5b918ad89989503fb957055431afa0af9fcbd4

Observation 12ffeb8f-f367-42bc-824b-ecf8f7633cd1 · inbound

Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation cites this paper.

Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T02:49:55.908142Z digest=sha256:bd84a8db65fea137e5450977fd98f88d1f4e49284a8a3299b28f49da73c261a0

Observation 7b12bf95-859e-466b-b5a2-da06b12f5c35 · inbound

Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models cites this paper.

Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:51:03.445109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T23:59:32.295839Z digest=sha256:a1fad630c976d31aba0088c38427365bfc91da5ead0d11ed1312e5a70d617170

Observation 9178735e-93d7-4cf0-baf9-de55e92fc93e · inbound

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing cites this paper.

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing Representation Engineering: A Top-Down Approach to AI Transparency

Reference 58

Resolution
malformed identifier
local_arxiv, observed 2026-05-11T22:16:25.475395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-08T03:09:43.879809Z digest=sha256:6aced66d9af4dbc7430f47a6c5b26a0cef1e51dfd34a49afbcd7a4f47a9d5d59

Observation 46e47c18-64ab-41a5-90b9-527975f3bab8 · inbound

Contextual Linear Activation Steering of Language Models cites this paper.

Contextual Linear Activation Steering of Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 41

Resolution
malformed identifier
local_arxiv, observed 2026-05-11T22:06:13.591263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-08T03:30:30.143805Z digest=sha256:0282546139f2867320e40496e63771f88c66da189a7e182f8f46dc283147c4fe

Observation 7db145e6-1bd2-472a-aeb0-edecf94c4434 · inbound

Architecture Determines Observability of Transformers cites this paper.

Architecture Determines Observability of Transformers Representation Engineering: A Top-Down Approach to AI Transparency

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:36:18.413775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-08T04:42:16.480831Z digest=sha256:53405a0880cd3f986e25f3a00c2d477b37e03ecca8856c117e9cac7a9bb1ee4b

Observation 134fbc5e-5576-4eda-ba27-bf4c1305da37 · inbound

Architecture Determines Observability of Transformers cites this paper.

Architecture Determines Observability of Transformers Representation Engineering: A Top-Down Approach to AI Transparency

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:32:29.229279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-13T07:32:09.488225Z digest=sha256:51b1b023c3521e827f53b4922764911070c835a41ab7a03625f2a393e92f0818

Observation 01c0fd95-92ca-49d5-9e57-56c9b8d4e1a1 · inbound

Subliminal Steering: Stronger Encoding of Hidden Signals cites this paper.

Subliminal Steering: Stronger Encoding of Hidden Signals Representation Engineering: A Top-Down Approach to AI Transparency

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:56:12.028836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-07T16:04:55.601088Z digest=sha256:85a15ee4df9cc7c491507cda49e5c8ab59eecc73432679aafc9b89bff45d7268

Observation 833b6a49-59ca-493d-a3e8-a849dd50e73d · inbound

MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks cites this paper.

MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:36:30.178588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-07T05:31:47.682478Z digest=sha256:4f354713a954fa4c390b0584bdd7cd14d733c9884845779c012fb6a7452fb697

Observation 6bc14e0c-4106-4efd-b57d-ce0c45c02748 · inbound

Attention Is Where You Attack cites this paper.

Attention Is Where You Attack Representation Engineering: A Top-Down Approach to AI Transparency

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:31:05.133859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-09T19:54:41.445447Z digest=sha256:79024c6828ad237f88fb9bfa88c00df34e0a56f9d72d0d4314579bc0509949ce

Observation daa7d426-a515-4056-bd55-ebbbeb8666ed · inbound

How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework cites this paper.

How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework Representation Engineering: A Top-Down Approach to AI Transparency

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:31:20.957306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T19:45:36.021741Z digest=sha256:73901eb66053a376f16ec48aff2eaa36ce0d02552e912f6cadc9314a363afa3a

Observation 8dba1b73-bf40-4353-9ea8-1c7c8ad533ea · inbound

Escaping Mode Collapse in LLM Generation via Geometric Regulation cites this paper.

Escaping Mode Collapse in LLM Generation via Geometric Regulation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:46:40.402651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-09T19:17:42.182356Z digest=sha256:1b098cec19c29584bface0e09f10a6dfc0384e60c36a49bc757fbd9e0433fd69

Observation b050556a-4f50-4973-a285-d929f3bfd2af · inbound

Escaping Mode Collapse in LLM Generation via Geometric Regulation cites this paper.

Escaping Mode Collapse in LLM Generation via Geometric Regulation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:05:30.553483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-07-01T08:02:42.741391Z digest=sha256:44c49a7ff2e923a56f3b53497583e614cb8a5a78e6db6c5535f9c17af3c545e6

Observation 13e6e8e2-7690-45e2-98cf-bc33a0cd8d83 · inbound

H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models cites this paper.

H-Probes: Extracting Hierarchical Structures From Latent Representations of Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T14:07:36.656164Z digest=sha256:1cf96210ccfb357779ac9e98a1ea74e02e1ff043265541e8f08df023bee6fa14

Observation 1c45c1ba-d450-49f4-9e5e-40d8325e778f · inbound

Latent Space Probing for Adult Content Detection in Video Generative Models cites this paper.

Latent Space Probing for Adult Content Detection in Video Generative Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:46:10.886489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-09T21:07:18.808339Z digest=sha256:dd8978983ac313f8348127eead904366eb62279acbc65e814654e2400bf3502d

Observation 884d8a7d-5ef8-4bed-afc3-533a22390db1 · inbound

Minimizing Collateral Damage in Activation Steering cites this paper.

Minimizing Collateral Damage in Activation Steering Representation Engineering: A Top-Down Approach to AI Transparency

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T18:58:30.056670Z digest=sha256:f5f5f92bf56ba1ae6e093be2ce6ef5b545a5bc8cf97d61ba50bd70db3353dedb

Observation 3021057e-6592-4be4-94cf-7ada47ea6b3b · inbound

A framework for analyzing concept representations in neural models cites this paper.

A framework for analyzing concept representations in neural models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:51:06.982163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T14:49:22.776209Z digest=sha256:76429ef0380e075d977a5e0c9fb0f97b4508a03b3a6f8bbd29842d623fba902d

Observation 12dfaaa3-b6b1-4b63-bfbf-5c7c671a4495 · inbound

The Cylindrical Representation Hypothesis for Language Model Steering cites this paper.

The Cylindrical Representation Hypothesis for Language Model Steering Representation Engineering: A Top-Down Approach to AI Transparency

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T00:15:09.263351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-01T00:10:29.122196Z digest=sha256:7a60811aeb87842a3999fd14ed29ca83778287ca50219d324b2c21e093fb4982

Observation 012db663-f289-46ed-8a92-6e07bb446206 · inbound

Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates cites this paper.

Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates Representation Engineering: A Top-Down Approach to AI Transparency

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-08T19:17:06.375875Z digest=sha256:9e50d7d6994f0852be61f0510668d57fbc88ea1405144f7ed6a2defb69c45126

Observation da824829-23a6-4e3f-814f-7171739e3393 · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T00:05:49.297190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:000613f78cfcef427a168fcf26ee55b18a16d89c3c6d744010ff56857f34db6f

Observation d085cf60-44c3-45da-82e3-047cb77b2aef · inbound

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs cites this paper.

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs Representation Engineering: A Top-Down Approach to AI Transparency

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:46:11.050355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-09T19:22:00.217729Z digest=sha256:a0f2a384560b75cc498c537741955674839531e639324ab9603f3256134639d4

Observation e2f42eac-0d05-44d2-b670-7afe29546a4f · inbound

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses cites this paper.

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses Representation Engineering: A Top-Down Approach to AI Transparency

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:46:31.258436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-07T02:16:49.785596Z digest=sha256:de79304d4601355b2b8fdfe11ca7e13de1c7907e9dcd2353357ce6bc1a408037

Observation c22e0a68-d59b-4126-a838-a63cc2079621 · inbound

Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis cites this paper.

Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis Representation Engineering: A Top-Down Approach to AI Transparency

Reference 6

Resolution
malformed identifier
local_arxiv, observed 2026-05-12T10:46:30.193230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-07T02:21:50.008687Z digest=sha256:0056716309f9803b0421487a20dffcf4393706eea2f2bd909115b8ff0f0618f3

Observation be1f7139-8e73-416a-9ac3-d303abf1eee3 · inbound

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It cites this paper.

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It Representation Engineering: A Top-Down Approach to AI Transparency

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:16:14.420457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-07T17:53:24.057211Z digest=sha256:aa9bda5acc4940e89a080321f1c6ae91ed7707a929bb2f3b8f54882f9c8ce0bd

Observation 5bdf8f21-e5df-4d01-a2e1-88b4c48fa818 · inbound

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It cites this paper.

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It Representation Engineering: A Top-Down Approach to AI Transparency

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.947124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-19T16:57:52.801399Z digest=sha256:13de58c9dd3423fe7475f93dd395bb23ad4de305335b5cb64ec26adf5df0034d

Observation 0295b8cd-b629-4c49-9868-6e34126261c9 · inbound

Steer Like the LLM: Activation Steering that Mimics Prompting cites this paper.

Steer Like the LLM: Activation Steering that Mimics Prompting Representation Engineering: A Top-Down Approach to AI Transparency

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-07T16:20:10.078995Z digest=sha256:cb84b182d7eb75da5b0759d5e84bfb56071077f63ed56df19e7c3d4ab63cec0b

Observation 730b57e2-b0e4-4332-b199-aee6585e82bc · inbound

Structural Instability of Feature Composition cites this paper.

Structural Instability of Feature Composition Representation Engineering: A Top-Down Approach to AI Transparency

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:11:41.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-10T06:40:00.507484Z digest=sha256:16168901ca23aa998d095423e018c7aa6fd25c4c45f5866c48d193cd09813cc9

Observation f2e29597-86eb-4e23-83ac-e356fb62eb68 · inbound

SLAM: Structural Linguistic Activation Marking for Language Models cites this paper.

SLAM: Structural Linguistic Activation Marking for Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:21:06.940740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-08T16:17:22.003460Z digest=sha256:cecca3a852660d70329a078c25cfd84ef7dfee7494f4c25bbf37c35181037c3b

Observation b211ae91-a9f5-4ac1-8868-7f556df831b2 · inbound

Negative Before Positive: Asymmetric Valence Processing in Large Language Models cites this paper.

Negative Before Positive: Asymmetric Valence Processing in Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:41:10.227342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-08T11:13:15.576248Z digest=sha256:221bcfbc391002784c2adedb1884e7e5d1c1587aa98fc80588b1ff34764261bb

Observation d72c1cac-57e1-4406-bd62-3f26d5e67ac3 · inbound

DataDignity: Training Data Attribution for Large Language Models cites this paper.

DataDignity: Training Data Attribution for Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T19:26:10.316198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-08T11:53:19.594779Z digest=sha256:9e894e598d0ac47907a151bb797be34889936b12febcf2bcc64e6e5749acf5a9

Observation 3c923213-ef65-42e3-a909-924b32ec3193 · inbound

On the Blessing of Pre-training in Weak-to-Strong Generalization cites this paper.

On the Blessing of Pre-training in Weak-to-Strong Generalization Representation Engineering: A Top-Down Approach to AI Transparency

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T18:36:08.753749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-08T14:59:19.883399Z digest=sha256:b6fc76227bf551ce9b11c8efc628b47e73d918654362e12a646c4d84b263cb73