Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T20:34:08.207848Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 94 inbound Pith citation observations for arXiv:2212.03827.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T20:34:08.207848Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:14:16.689102Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
44 of 44 outbound references displayed
External citation measurements
45
pith, observed 2026-08-05T02:28:24.338817Z
Observation 7054bd77-617c-4b80-8bfd-00518389ec30 · outbound
Discovering Latent Knowledge in Language Models Without Supervision A General Language Assistant as a Laboratory for Alignment
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 33d77505-15b6-4fe6-93ce-2a5dfb2592dd · outbound
Discovering Latent Knowledge in Language Models Without Supervision Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7b72cc2-45c0-47b2-95a7-a421c5a34e29 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 41f2e11f-2b54-4efb-abe6-81569996d8c1 · outbound
Discovering Latent Knowledge in Language Models Without Supervision On the Opportunities and Risks of Foundation Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c402b99c-6335-44fc-b359-d92a8708faba · outbound
Discovering Latent Knowledge in Language Models Without Supervision Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 316431e8-35a7-4f2f-b5c6-35f228550c4d · outbound
Discovering Latent Knowledge in Language Models Without Supervision Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc95caa8-bb73-4a7e-b699-88e094c64a50 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Supervising strong learners by amplifying weak experts
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5698c28a-791c-4963-9f57-c48273c6b764 · outbound
Discovering Latent Knowledge in Language Models Without Supervision BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 12269bc7-6dd8-4ba7-89e7-e45b05c82c0f · outbound
Discovering Latent Knowledge in Language Models Without Supervision Truthful AI: Developing and governing AI that does not lie
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 560aaf69-397e-4a37-9c29-567103800a68 · outbound
Discovering Latent Knowledge in Language Models Without Supervision DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e200fab-3720-44fa-80e5-b0ff20d7eda9 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Unsolved Problems in ML Safety
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 548d9f5b-f722-487f-8316-505edf7f8b12 · outbound
Discovering Latent Knowledge in Language Models Without Supervision AI safety via debate
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e76b8ba6-ed1d-4f2c-a0a5-65d7d0c74ac0 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d0cd27f-e5b7-485f-ae64-cf47f5e8f508 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Alignment of Language Agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c03ffb76-70c4-487c-8b84-28b58ab7dc65 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Ground-Truth Labels Matter: A Deeper Look into Input-Label Demonstrations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f03c4e3a-47fe-4779-ad06-5deff716a1fa · outbound
Discovering Latent Knowledge in Language Models Without Supervision Adam: A Method for Stochastic Optimization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 838a7ba4-b93f-4f1f-9f8e-aa2f772f759b · outbound
Discovering Latent Knowledge in Language Models Without Supervision Scalable agent alignment via reward modeling: a research direction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac205680-71ff-44e0-804c-a37f7454ead5 · outbound
Discovering Latent Knowledge in Language Models Without Supervision RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 475dc91b-c9e4-42ba-b257-490c63ef2b7c · outbound
Discovering Latent Knowledge in Language Models Without Supervision Decoupled Weight Decay Regularization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7e7ce69-a935-47ec-8949-7b84c00a5a3a · outbound
Discovering Latent Knowledge in Language Models Without Supervision On Faithfulness and Factuality in Abstractive Summarization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3a61323-0f87-4ad2-881d-0289e4a60ec9 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Teaching language models to support answers with verified quotes
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b18b2f4-47b0-4b94-b739-f752e4157c55 · outbound
Discovering Latent Knowledge in Language Models Without Supervision MetaICL: Learning to Learn In Context
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1583a627-e650-41c0-b44b-1d0cf97e21a8 · outbound
Discovering Latent Knowledge in Language Models Without Supervision WebGPT: Browser-assisted question-answering with human feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c95bebdc-9c95-457b-b97a-ebd6e96d1c0d · outbound
Discovering Latent Knowledge in Language Models Without Supervision Training language models to follow instructions with human feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 326f6614-903a-48fe-96a1-774af62f3fac · outbound
Discovering Latent Knowledge in Language Models Without Supervision Red Teaming Language Models with Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 96809c38-c7d3-495c-b8ae-2ee028c92ed7 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d69ab1c6-9bd7-46b7-9bea-97c93036c77b · outbound
Discovering Latent Knowledge in Language Models Without Supervision SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e07a74a8-8a1d-4c2e-8c5d-5556d2523c75 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29f5c200-1097-4bf6-bd93-ed2299d2330e · outbound
Discovering Latent Knowledge in Language Models Without Supervision Multitask Prompted Training Enables Zero-Shot Task Generalization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a21247f0-bcbd-4776-84ee-cd02e05a69b1 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Learning to summarize from human feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b05242b8-3f58-4cb6-b355-12ec5e3ff3aa · outbound
Discovering Latent Knowledge in Language Models Without Supervision FEVER: a large-scale dataset for Fact Extraction and VERification
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dac2e12b-a208-4dac-be6e-f423cf1e2fe6 · outbound
Discovering Latent Knowledge in Language Models Without Supervision GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23635a30-246e-4ba9-b4e8-79ce61bc5709 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Finetuned Language Models Are Zero-Shot Learners
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b5fc1b9-5018-4f6b-9899-270224ad3412 · outbound
Discovering Latent Knowledge in Language Models Without Supervision HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44ec5d43-9a94-4655-9930-48e61a3a461d · outbound
Discovering Latent Knowledge in Language Models Without Supervision Calibrate Before Use: Improving Few-Shot Performance of Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bab857f6-9b58-4109-b43c-69fa69990089 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Prompt Consistency for Zero-Shot Task Generalization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 93f9bf49-bd74-4324-a474-56a74bb934ea · outbound
Discovering Latent Knowledge in Language Models Without Supervision Is 2+2=4? Yes
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44ae3498-4e2b-4f8d-b4be-4a3b20ebf014 · outbound
Discovering Latent Knowledge in Language Models Without Supervision hidden states
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 669c4cac-f670-41fb-99f1-d862ba1ceac1 · outbound
Discovering Latent Knowledge in Language Models Without Supervision [text] = I loved this movie
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7e8f239f-3405-4df5-898c-b2342791e464 · outbound
Discovering Latent Knowledge in Language Models Without Supervision ‘ [content]
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee62ca3e-f61e-4b6d-a780-8e1243e6fc66 · outbound
Discovering Latent Knowledge in Language Models Without Supervision Here the label is a short sentence
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7abd8d6-dcc9-42a3-b49b-3cc077d35637 · outbound
Discovering Latent Knowledge in Language Models Without Supervision ‘ [premise]
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b215f2e-1d43-44a0-bf79-6af2a0aa907d · outbound
Discovering Latent Knowledge in Language Models Without Supervision [label]” is “negative
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bf9e264-3cfc-4706-9089-47164b0a879e · outbound
Discovering Latent Knowledge in Language Models Without Supervision yes” or “no
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 972a7f42-7d52-416a-9f5e-4c1a96b903d4 · inbound
The Internal State of an LLM Knows When It's Lying Discovering Latent Knowledge in Language Models Without Supervision
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a436c8ed-cd9c-44ce-ac8c-599dc9710c33 · inbound
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions Discovering Latent Knowledge in Language Models Without Supervision
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a4a42f7-f3a3-4b76-a8e5-d13a63289cfc · inbound
Refusal in Language Models Is Mediated by a Single Direction Discovering Latent Knowledge in Language Models Without Supervision
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a856920c-44b4-4c1e-ac61-6f54e23e35f8 · inbound
Training Language Models to Self-Correct via Reinforcement Learning Discovering Latent Knowledge in Language Models Without Supervision
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 99afcef3-01b2-451f-82c4-4b463a8691ba · inbound
Mechanistic Interpretability Needs Philosophy Discovering Latent Knowledge in Language Models Without Supervision
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a1a899e-d56f-4157-9daa-0fef7608adf1 · inbound
NPO: Learning Alignment and Meta-Alignment through Structured Human Feedback Discovering Latent Knowledge in Language Models Without Supervision
Reference 2000
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42cad935-733f-447c-b499-5832851180d8 · inbound
LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788a296d-7b76-4c98-8253-351e8993630b · inbound
A Survey on Data Security in Large Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208a2103-5324-43ae-af81-1de597b8beb4 · inbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Discovering Latent Knowledge in Language Models Without Supervision
Reference 224
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34721bb2-9705-4205-8fb2-c98b1cc3fd4e · inbound
Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d35989-fd01-498f-bf30-5a86313d7fc2 · inbound
STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment Discovering Latent Knowledge in Language Models Without Supervision
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6239db5-d241-4e32-b4db-0f8ff4549ecf · inbound
Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db747075-a0cc-407d-971f-457774cc3c43 · inbound
Can LLMs Lie? Investigation beyond Hallucination Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0965d9b4-8212-4343-8f0a-c922e0bfc76c · inbound
Cross-Layer Attention Probing for Fine-Grained Hallucination Detection Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 737f2084-1b16-47e0-85e0-155f7b7506aa · inbound
Unsupervised Hallucination Detection by Inspecting Reasoning Processes Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92854d8-3d78-4b48-9f3d-f77b14696bae · inbound
HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling Discovering Latent Knowledge in Language Models Without Supervision
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9f8d589-fb1d-421f-b447-01e879dbc9e5 · inbound
Neural Message-Passing on Attention Graphs for Hallucination Detection Discovering Latent Knowledge in Language Models Without Supervision
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3ccaa2-bfbe-4d8b-b28e-3105ba86be42 · inbound
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning Discovering Latent Knowledge in Language Models Without Supervision
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2de349-72c6-4f7b-ab80-61ce05f197a7 · inbound
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd69cf10-c629-4061-80ef-e2e15dd1ebc3 · inbound
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bded4634-bc39-47e4-84ff-415f09055973 · inbound
Emergent Manifold Separability during Reasoning in Large Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4aace6aa-ea9d-41de-b8b7-1bdd9cb7ca87 · inbound
Prompt Injection as Role Confusion Discovering Latent Knowledge in Language Models Without Supervision
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d64fab97-26cc-4a83-84bc-4d05345c50ef · inbound
How do LLMs Compute Verbal Confidence Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7be395c5-0fd1-4d93-a8ee-072d92950f40 · inbound
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b14138c-2eaf-4c4d-8303-afbd55836c96 · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2834ec8-8899-4835-81b0-70fbcf06de46 · inbound
Weakly Supervised Distillation of Hallucination Signals into Transformer Representations Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5494bc89-f8f3-40da-974c-95101eb4dedd · inbound
The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 13e3a36c-5c8b-49ba-9e03-63963fc621b9 · inbound
Learning Uncertainty from Sequential Internal Dispersion in Large Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 81c5f600-eaa4-4f86-b2a1-6cd0cd3819ec · inbound
How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them Discovering Latent Knowledge in Language Models Without Supervision
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 794bba6f-e38f-49ce-99b6-7b1ae2127bba · inbound
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders Discovering Latent Knowledge in Language Models Without Supervision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 20783210-ed4f-436e-a807-f47148b73a8c · inbound
How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51eddca8-d954-4d0a-92ae-a5689df686a0 · inbound
Perturbation Probing: A Two-Pass-per-Prompt Diagnostic for FFN Behavioral Circuits in Aligned LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e77e8cdf-ab76-4031-8f23-5d13648ba686 · inbound
Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability Discovering Latent Knowledge in Language Models Without Supervision
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf765001-42ef-4f78-92ff-c15d497f1ed4 · inbound
Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes Discovering Latent Knowledge in Language Models Without Supervision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 240a3d52-49f6-40e1-87b3-66b0bce81b0c · inbound
The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations Discovering Latent Knowledge in Language Models Without Supervision
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0be2a14b-9cd2-4eef-9ae3-3db49ac3544f · inbound
Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs Discovering Latent Knowledge in Language Models Without Supervision
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 148004bd-cf73-4510-ae50-056f92d70652 · inbound
Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs Discovering Latent Knowledge in Language Models Without Supervision
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19837d01-1976-41dd-b6a2-05f2d0360a41 · inbound
LLM Agents Already Know When to Call Tools -- Even Without Reasoning Discovering Latent Knowledge in Language Models Without Supervision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8e04aa4-4e78-48dc-ac11-00e62a8cd347 · inbound
LLM Agents Already Know When to Call Tools -- Even Without Reasoning Discovering Latent Knowledge in Language Models Without Supervision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f003387-f22f-4da1-a30d-ae18832fadc6 · inbound
Positive Alignment: Artificial Intelligence for Human Flourishing Discovering Latent Knowledge in Language Models Without Supervision
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68f47cbd-83af-403b-9446-21e8ef19ac46 · inbound
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Discovering Latent Knowledge in Language Models Without Supervision
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d235f8c1-d6db-4105-b730-660c8b006b2c · inbound
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Discovering Latent Knowledge in Language Models Without Supervision
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f7b26cd2-f8f1-443c-bc01-60f5f5733347 · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Discovering Latent Knowledge in Language Models Without Supervision
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b12cf2a5-3e99-4956-ab39-40a3e91f015a · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Discovering Latent Knowledge in Language Models Without Supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8445a34d-820e-40ce-af87-ef1254df4472 · inbound
Trust or Abstain? A Self-Aware RAG Approach Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation abc1a637-f32f-4bf0-965d-eb101dd82cdd · inbound
Manifold-Guided Attention Steering Discovering Latent Knowledge in Language Models Without Supervision
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9a0d9cb4-ca11-4c2e-8336-fb1ffae02cef · inbound
Reading Calibrated Uncertainty from Language Model Trajectories Discovering Latent Knowledge in Language Models Without Supervision
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ebe266b5-2bd6-4f20-bde9-98b6b079f2a8 · inbound
Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a37d5119-93fe-4af7-b9a2-f7cd5203f68c · inbound
Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking Discovering Latent Knowledge in Language Models Without Supervision
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d5845af-5a7e-4218-8c30-489698a744e1 · inbound
The Attentional White Bear Effect in Transformer Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 425ad979-64c6-4b4e-bfec-1a9a76639028 · inbound
Can LLMs Use Linguistic Uncertainty Markers to Reliably Reflect Intrinsic Confidence? Discovering Latent Knowledge in Language Models Without Supervision
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea620236-5c89-40f9-9a39-07223841e287 · inbound
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet Discovering Latent Knowledge in Language Models Without Supervision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 032fde70-6b90-4177-963a-7df421816826 · inbound
UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68d539dc-ec8c-445b-a156-a838d7880a14 · inbound
CANARY: Zero-Label Detection of Fine-Tuning Contamination in Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation edcd55e5-5802-4b5e-8faa-c1aea132a358 · inbound
Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8c7a4ce7-7193-4630-ab35-d123cb01ad32 · inbound
Consistency Training Can Entrench Misalignment Discovering Latent Knowledge in Language Models Without Supervision
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b70fcb39-61c6-43c1-8e5a-8a6447d5c9fa · inbound
CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 726e4a78-ae6f-4130-850d-d33cb3276854 · inbound
LLM Self-Recognition: Steering and Retrieving Activation Signatures Discovering Latent Knowledge in Language Models Without Supervision
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f1590fe2-9c58-4a74-80cb-8fb82e85583c · inbound
Adversarial Robustness of Activation Steering in Large Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d2c0df88-dd3e-4fd3-a531-d1574f04206c · inbound
Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fe7d966-cb83-4746-a98c-f67677d0a77f · inbound
PRISM: Recovering Instruction Sets from Language Model Activations Discovering Latent Knowledge in Language Models Without Supervision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d2db5217-5ae7-4d8a-98da-78e0d5e50ee5 · inbound
Toward Calibrated, Fair, and accurate Deepfake Detection Discovering Latent Knowledge in Language Models Without Supervision
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29b05033-4b28-41f9-8174-14e58829eee9 · inbound
Forecasting Future Behavior as a Learning Task Discovering Latent Knowledge in Language Models Without Supervision
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f33b6223-44aa-466e-92ca-5aa5c8c7b925 · inbound
Pre-Generation Hallucination Detection in Large Language Models via Soft-Target Attention Probing Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7f67689c-c851-4bea-ab5b-dab52f299521 · inbound
When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a6bb1225-a0a3-4c8e-a376-655bf79aa0e8 · inbound
Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents Discovering Latent Knowledge in Language Models Without Supervision
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1cdbb5a5-2feb-45cb-8bf8-99aafe702d90 · inbound
Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0cfb732d-67cc-4985-a2fc-187b0d3880c8 · inbound
Radical AI Interpretability Discovering Latent Knowledge in Language Models Without Supervision
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 815fb3b0-c0b6-4a99-afab-60022749c5a6 · inbound
Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions Discovering Latent Knowledge in Language Models Without Supervision
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c8c957e-203c-4a69-89b5-636ecf12beba · inbound
From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 988d2075-4bd9-45aa-902e-1ed74210f20c · inbound
The strength of clinical evidence is recoverable from language model representations but not from their stated grades Discovering Latent Knowledge in Language Models Without Supervision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b3528d69-d2c3-49fb-8a45-7bc3bad167a1 · inbound
Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring Discovering Latent Knowledge in Language Models Without Supervision
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51c3f320-af26-410a-b377-1e688649833e · inbound
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65575cb8-6020-4fc1-b1c3-ec9f82a82757 · inbound
PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption Discovering Latent Knowledge in Language Models Without Supervision
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac1e7194-fc01-4a14-bd82-fbc71b4e4c05 · inbound
Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e99762da-d12e-4ea4-b327-88f774e91423 · inbound
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc5b659a-cfca-421e-948d-a9b4b4bd793c · inbound
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb5c000-5b4c-40a3-b989-4f544d912bc1 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Discovering Latent Knowledge in Language Models Without Supervision
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ddb4ac6-1bdc-4ba8-918d-16f06d406da1 · inbound
Weak-to-Strong Generalization via Direct On-Policy Distillation Discovering Latent Knowledge in Language Models Without Supervision
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b1ff89-6bd2-4fc3-9ea7-260b4678a84c · inbound
Dissociating the Internal Representations of Sycophancy in LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9bbff01e-d216-4a27-9234-b4bd2691c602 · inbound
Dissociating the Internal Representations of Sycophancy in LLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef966a6b-d11d-4927-a5ff-07acb8c4ad06 · inbound
Prompt Compression via Activation Aggregation Discovering Latent Knowledge in Language Models Without Supervision
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d105eec-e4fb-4170-aeb5-ce3b335bcbf2 · inbound
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Discovering Latent Knowledge in Language Models Without Supervision
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153b3803-5472-4faf-b276-4b1bb31ccf6d · inbound
The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs Discovering Latent Knowledge in Language Models Without Supervision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19d968ca-3e05-41c8-a177-2d7064d74d55 · inbound
Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States Discovering Latent Knowledge in Language Models Without Supervision
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 029b3bbe-157e-4230-9e2a-02fd015bb33e · inbound
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias Discovering Latent Knowledge in Language Models Without Supervision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d54499-1682-4ac0-92b3-a092556fbc40 · inbound
The Computational Basis of Confidence in Large Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc4cf203-9333-48e0-bd82-110f326d7d03 · inbound
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't Discovering Latent Knowledge in Language Models Without Supervision
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d3e08a-91a9-4b84-a73b-1f135effb8c3 · inbound
Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering Discovering Latent Knowledge in Language Models Without Supervision
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddf48a2-c533-4b2d-accc-03a15f001caa · inbound
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making Discovering Latent Knowledge in Language Models Without Supervision
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bcafe71-c745-4368-8b24-f0556c70d460 · inbound
The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models Discovering Latent Knowledge in Language Models Without Supervision
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca63d1e-16b8-453c-8966-7f98fb4a1a71 · inbound
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes Discovering Latent Knowledge in Language Models Without Supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31fcd1be-d249-4209-9f92-08c7bd29507f · inbound
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model Discovering Latent Knowledge in Language Models Without Supervision
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac76bc1-07cc-4137-a403-25d788855de2 · inbound
Securing Multimodal AI through Internal Information Decomposition Discovering Latent Knowledge in Language Models Without Supervision
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.