Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T04:33:54.058475Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2604.18519.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T04:33:54.058475Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T21:28:25.577739Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T21:35:04.761308Z
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d5b2a56b-da7c-4d7e-99ac-83fb41c9db39 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca5b829a-b70c-420a-ab51-3c147be53ba6 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the AAAI conference on artificial intelligence , volume=
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d61a91f-2d50-47b3-a9f9-22c53dd92841 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef8cd7b4-e8c6-4fcc-a861-7dcac8db4b44 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e4d273c-0bd7-403f-bafe-ec9e1246d42a · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f57f7dc-df1c-4f30-927d-2128b3986f66 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d00c611a-d5db-4180-95cc-80bc559621db · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1116b0a4-0e5d-4143-9f40-f09da8eb3e57 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 448b001b-4065-4a18-a656-900e9c4bfb9f · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6efb3c7b-b79d-4f91-a7e1-05a7c2bb9202 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0afad29-862c-4db9-aa1b-a349345a1af3 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations LLMs Encode Harmfulness and Refusal Separately
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea42c941-7525-4501-9450-11bffdbb0f0b · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7f1df501-f39a-4bbc-937e-38143a422378 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Qwen3Guard Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84a353e3-dee7-49ab-9b6a-83160348045d · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7645e6d1-bd18-4654-b0f6-c8e614faa764 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc88ce6c-433a-4c59-9e53-59accdf3f005 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in neural information processing systems , volume=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60ae356d-b6f3-4147-ae6d-213293efd60f · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining , pages=
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 523891dd-59c8-4857-a9b1-cfa5e10863a8 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Findings of the Association for Computational Linguistics: ACL 2024 , pages=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ab56c69-546a-4139-a7be-3b6b496fbf07 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations The Eleventh International Conference on Learning Representations , year=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bd26b10-29c1-4ccf-a2be-4546be3cd448 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a8ce3e2-c714-4813-9926-af1fd2a9d3b9 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations BERT Rediscovers the Classical NLP Pipeline
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b18c8fae-7c8a-41f2-8f99-48f1ad45b8c4 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43d98db9-d6d8-498c-bf60-daebb467c144 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations The information bottleneck method
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e34d1d47-aa1f-4d77-8861-8fe651e41ff1 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations 2022 , journal=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69f7e846-7b42-45c1-a132-6da1787dae30 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Transactions on Machine Learning Research , year=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69121aba-0375-44d3-ba70-9283af3100a9 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 42bdf367-d710-42e8-8428-dcbef79212a7 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2022 conference on empirical methods in natural language processing , pages=
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b69e8a9-a4ab-4deb-abb0-dd3659881546 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4b800ca-a3da-43b0-a7b1-a1fda654f7fb · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Layer by Layer: Uncovering Hidden Representations in Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0184305d-8bb6-4eff-83ea-c80f32d7ecf2 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f26b44fe-ecee-4a67-9ae5-c4390c5e1a6a · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Organization of Knowledge and Advanced Technologies
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fec285d-6f79-4c4b-af1c-b823e1b66dcf · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Efficient LLM Moderation with Multi-Layer Latent Prototypes
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d09d643a-6e56-4864-a8ff-6dafdace5ff8 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Safety Layers in Aligned Large Language Models: The Key to LLM Security
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c21951ba-222c-4ee3-8d29-9d89bd632016 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 288a5fb8-4d8c-41d7-add8-f2f1fab2587d · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv preprint arXiv:2510.06594 , year=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1935def0-d11b-4f5c-8a33-228208403b18 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Linearity of Relation Decoding in Transformer Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da6ffd10-3a93-4014-835a-f3291df9f26e · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Eliciting Latent Predictions from Transformers with the Tuned Lens
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a266207-ff1e-4b3c-a046-ba8b81cf6fb0 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Companion Proceedings of the ACM on Web Conference 2025 , pages=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a5eb6c6-89f2-4c7a-ae93-07f870c29af0 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Findings of the Association for Computational Linguistics: ACL 2025 , pages=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95ba164b-e9b8-4273-83dc-726ed179b39c · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations The Linear Representation Hypothesis and the Geometry of Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa7e2d14-1c43-4d50-8740-2df308cfcdac · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Understanding intermediate layers using linear classifier probes , url =
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af32c80a-6a32-45b6-a975-6977ccb54369 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations An introduction to variable and feature selection , url =
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec5f792e-2537-4fb4-a19e-bd90e7521df7 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Qwen3 Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa71f58b-386e-47f6-9ee3-b2b19877f419 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv e-prints , pages=
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b434071-60ab-45dd-94a1-f0891d6c830a · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Lightweight Safety Classification Using Pruned Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37f1d075-d1a1-44e1-bf1c-d313c220631a · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations ShieldGemma: Generative AI Content Moderation Based on Gemma
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed317d6a-02e1-4c85-9ecf-89b825591580 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Generative or Discriminative? Revisiting Text Classification in the Era of Transformers
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4dd6b657-03f0-4045-9038-4df151e4383b · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations The Thirteenth International Conference on Learning Representations , year=
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e0f907f-bdd2-42f2-8de5-68a3b7994bf9 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 134b24df-8dfe-496c-a19f-abc97a7ff2f8 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations IEEE Robotics and Automation Letters , volume=
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dff3646d-ff43-490f-af8f-75e286bb77b7 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations RDI: An adversarial robustness evaluation metric for deep neural networks based on model statistical features
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0352f629-72d9-4450-8d61-fe393b41a693 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25396c11-0073-4387-9657-1dcad9e06309 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Scaling laws for neural language models , url =
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ecedd5a7-16f4-4a84-b195-84538c5d991d · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b56e116-1bfb-49f2-85bf-d149403d1cdf · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations author=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f437287a-2719-4f3a-9179-e74fe41ecda4 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c8b0a62-c9a2-44f9-b6b3-d4a9fbeb5989 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations 2025 , howpublished =
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e552a06-6d22-4761-a2d0-2b53260aa24c · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations 2025 , howpublished =
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b9dbb1e-d773-4a52-ac22-4b31eac9f159 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ffef57a-39c4-491a-8344-0b79e4db3752 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7426ea55-99ba-44c8-a9e5-397ca6e2f056 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96928c83-816f-4f73-adf8-bb0302d28415 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 190f9fcb-274b-4fad-91b2-00244340b913 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv preprint arXiv:2508.03550 , year=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c401dee2-40a4-4c16-a905-26cd98ef813d · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the IEEE conference on computer vision and pattern recognition , pages=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd9f2d22-03f3-4eac-a64c-2d087444a635 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b760b1f-10fb-4087-a3cf-f513f2d6604f · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations 2019 , eprint=
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0c1504c-2ded-4914-9a9f-f6a8b85423c4 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Hate speech detection and racial bias mitigation in social media based on bert model.PLOS ONE, 15(8):1–26, 08 2020
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc4b602a-2457-4275-b6ea-74d178d22fdf · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations 2021 , isbn =
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10d91182-2d94-4175-8282-2944ecafb995 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations H ate BERT : Retraining BERT for Abusive Language Detection in E nglish
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02b07d0d-af88-48fd-8079-717a37ffa34a · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations 2024 , eprint=
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06a6ec07-824a-4b46-8aa8-5c30d572501c · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations and Tay, Yi and Sorensen, Jeffrey and Gupta, Jai and Metzler, Donald and Vasserman, Lucy , title =
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5abbd1ea-32e1-4861-b346-9b5e7c58d9c5 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations 2025 , eprint=
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0421a101-051c-442a-a6b4-4b3a122c85d1 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e9acd6d-91b0-4870-a830-94de92d2b6cd · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a853100-0621-4b91-8431-916beeb54807 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Linear Representations of Sentiment in Large Language Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36b3f119-44c3-4de0-8410-ade08be64a1a · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 279548d1-4dae-4d91-9093-32759b537cbf · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 28th ACM international conference on information and knowledge management , pages=
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2aa0f24-f7bf-4c7f-9eb9-2c042ce9d75b · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv preprint arXiv:2510.18081 , year=
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b0a7544-a00d-4285-9474-24c935d534b5 · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Curvalid: Geometrically-guided adversarial prompt detection
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79fa70d3-2184-45cb-b67e-8ed50c7fd44f · outbound
LLM Safety From Within: Detecting Harmful Content with Internal Representations Neural networks , volume=
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85e0dbd5-f316-43eb-a156-6af8652f9f7f · inbound
MINER: Mining Multimodal Internal Representation for Efficient Retrieval LLM Safety From Within: Detecting Harmful Content with Internal Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89bd88d0-aa8e-4cca-80eb-23bae07202bb · inbound
LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems LLM Safety From Within: Detecting Harmful Content with Internal Representations
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07051512-4efa-4e18-b1c7-73f64e217c15 · inbound
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue LLM Safety From Within: Detecting Harmful Content with Internal Representations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.