Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2503.10965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-15T05:51:54.267975Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 1f8e93d7-8c25-406e-ad64-0b7e8daba457 · inbound
Internal Deployment in the AI Act Auditing language models for hidden objectives
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e2285006-db6a-4264-814f-b9b95cfdf594 · inbound
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves? Auditing language models for hidden objectives
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation ecabf288-40ff-4f20-8eee-4284d200fcb7 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Auditing language models for hidden objectives
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7919d241-d8ef-400d-a631-0f3936192d04 · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Auditing language models for hidden objectives
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5ce426cb-d4e5-4ef4-960a-c6d84b54944b · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Auditing language models for hidden objectives
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 56fdb4a2-ef28-4146-ba4a-8e4e09a98f61 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 12f71361-56e3-40e4-b345-1c4336b2f5b3 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d3c0b841-123a-4bde-a163-775c7ade017e · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 41e33e97-cf37-43ac-b5a3-0e315219a3de · inbound
Positive Alignment: Artificial Intelligence for Human Flourishing Auditing language models for hidden objectives
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f9e9e36c-bd9e-4583-b4c1-7b602369d35d · inbound
Deep Minds and Shallow Probes Auditing language models for hidden objectives
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3aa6eb90-0c30-46f6-9c73-13884c2f24f5 · inbound
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Auditing language models for hidden objectives
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5e6652c7-54d7-4f01-a9f2-a4d661157ebc · inbound
Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands Auditing language models for hidden objectives
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3550bd84-1b41-4078-bf8e-2a075683e031 · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Auditing language models for hidden objectives
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 819b93fe-de95-498a-9cb6-748d9e3e96e5 · inbound
PRISM: Recovering Instruction Sets from Language Model Activations Auditing language models for hidden objectives
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6c3f657a-63bc-48ee-90d3-6563d1ff96a7 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Auditing language models for hidden objectives
Reference 250
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4eadd4bd-af71-4321-9a60-26aa20c3d2ec · inbound
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms Auditing language models for hidden objectives
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7be5d1e9-cb3f-4c8b-a787-ba932cf5e064 · inbound
RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue Auditing language models for hidden objectives
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 093f88b0-768b-456a-a7c9-f556eb286e15 · inbound
Self-CTRL: Self-Consistency Training with Reinforcement Learning Auditing language models for hidden objectives
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1444b25b-1bd5-4a3d-b6ea-21c396c7ad16 · inbound
Channel Location Constrains the Auditability of Subliminal Learning Auditing language models for hidden objectives
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f52bdd47-9f53-4350-a462-8c2dd4ed46d3 · inbound
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology Auditing language models for hidden objectives
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3b2d4f8e-a49e-47ed-9b54-1dc75367f11d · inbound
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric Auditing language models for hidden objectives
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.