Pith. sign in

Paper Citation Record · LEDGER

LLM Safety From Within: Detecting Harmful Content with Internal Representations

As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2604.18519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18519 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T04:33:54.058475Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:28:25.577739Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T21:35:04.761308Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact12
  • verified fuzzy36
  • unresolved1
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch30

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5b2a56b-da7c-4d7e-99ac-83fb41c9db39 · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

LLM Safety From Within: Detecting Harmful Content with Internal Representations ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.136237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:ca78d0ca39b397c7ec35b219093d48b754ab36b70f1a918bca0d081878a2bfd9

Observation ca5b829a-b70c-420a-ab51-3c147be53ba6 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.338928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:a658f9ae0cccef3ce614f1507e80f1f5592c9592862fe6dd5f3497a1b623d48f

Observation 5d61a91f-2d50-47b3-a9f9-22c53dd92841 · outbound

This paper cites AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts.

LLM Safety From Within: Detecting Harmful Content with Internal Representations AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.127272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:da89a88567617bd0fcf6ae8963316780821ca5ddf7df8ad85c70c14445bbd12c

Observation ef8cd7b4-e8c6-4fcc-a861-7dcac8db4b44 · outbound

This paper cites Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.119498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:d925c8548bc746cf0970f2dd0941c4db453edd00e6c23e05a6680245af3f6374

Observation 3e4d273c-0bd7-403f-bafe-ec9e1246d42a · outbound

This paper cites SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.111469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:2082612ad952339c3dc7bbf853413a4b60e197de414229e07f91b3d6578c31f0

Observation 3f57f7dc-df1c-4f30-927d-2128b3986f66 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

LLM Safety From Within: Detecting Harmful Content with Internal Representations HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:06:33.939395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:05075d728601638fb5b5f098892a914cd06943734af8f343cc1843811150d863

Observation d00c611a-d5db-4180-95cc-80bc559621db · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.331178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:075f50afb898efc47d18a2e5ae4ec210e4cdd805218664845762025536fe3497

Observation 1116b0a4-0e5d-4143-9f40-f09da8eb3e57 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

LLM Safety From Within: Detecting Harmful Content with Internal Representations PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.122827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:f897a59671ce151ec5e61eca67d92846dc4f71441baeeafc1d903e8ba86d3841

Observation 448b001b-4065-4a18-a656-900e9c4bfb9f · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.335141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:24cbd895f4602160771b5bed892bff0752ae996c096b13e5ec45186f52694dc9

Observation 6efb3c7b-b79d-4f91-a7e1-05a7c2bb9202 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:51:50.894684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:7ef607673c8171325f3720c158ceda342869cfca14910dd03174bcfa332b93db

Observation e0afad29-862c-4db9-aa1b-a349345a1af3 · outbound

This paper cites LLMs Encode Harmfulness and Refusal Separately.

LLM Safety From Within: Detecting Harmful Content with Internal Representations LLMs Encode Harmfulness and Refusal Separately

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:18:08.851986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:b3a2056e3844752b871292a0085d8a9be75f5764d9a1712b01591226735e56da

Observation ea42c941-7525-4501-9450-11bffdbb0f0b · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.319477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:fb8e17c21d2a6a83650a0daf19a750aba7696db45f6631f2c5b4f505fa936fe2

Observation 7f1df501-f39a-4bbc-937e-38143a422378 · outbound

This paper cites Qwen3Guard Technical Report.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Qwen3Guard Technical Report

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:33:37.835217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:5c3b9f44c3a2c6ae837cba7578e6ccc1863ac19d265af4d969ba6b07fb379bcd

Observation 84a353e3-dee7-49ab-9b6a-83160348045d · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.922788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:67f67abf5c1c1a443c15c2dd4fb6af8230111cd21fd7703c3a0e7f298afecde8

Observation 7645e6d1-bd18-4654-b0f6-c8e614faa764 · outbound

This paper cites Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.143551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:f2b76288319f1231df84c27e86b4a255427f4b0079c3c35c8b375430bb049080

Observation cc88ce6c-433a-4c59-9e53-59accdf3f005 · outbound

This paper cites Advances in neural information processing systems , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in neural information processing systems , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.323469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:a9008a3e069dc2e6aea004d9fd72bbb2b971db8b12c1ed82336df9b4f1904265

Observation 60ae356d-b6f3-4147-ae6d-213293efd60f · outbound

This paper cites Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining , pages=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.327399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:6acab7f3bf132b0950e8025f72fdc461e4011c4e8a093374b6b1783fec63e0b2

Observation 523891dd-59c8-4857-a9b1-cfa5e10863a8 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.978394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:c27ff762bb9f1312074b2d423a772dc06652a8b78489703707687084a7d272c2

Observation 7ab56c69-546a-4139-a7be-3b6b496fbf07 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations The Eleventh International Conference on Learning Representations , year=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.008478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:1080a6660f5bf8faf1c8e204003952e6bdc66f570e9d0fa84f7dc700890dfdb5

Observation 1bd26b10-29c1-4ccf-a2be-4546be3cd448 · outbound

This paper cites Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.011054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:cbe7d16d018699e78b52638dbcba4d51a7831dc8d1d951928c07e9379b99f489

Observation 5a8ce3e2-c714-4813-9926-af1fd2a9d3b9 · outbound

This paper cites BERT Rediscovers the Classical NLP Pipeline.

LLM Safety From Within: Detecting Harmful Content with Internal Representations BERT Rediscovers the Classical NLP Pipeline

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.051467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:2d760febce0b7a3df980e39b41e635263fd5be2b574f77a906ed1b4959204bff

Observation b18c8fae-7c8a-41f2-8f99-48f1ad45b8c4 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.932781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:88c4c81ff63fcf00fd2a166e4a90e65be1f6bf2e1b45cee94050c2b31e3a0695

Observation 43d98db9-d6d8-498c-bf60-daebb467c144 · outbound

This paper cites The information bottleneck method.

LLM Safety From Within: Detecting Harmful Content with Internal Representations The information bottleneck method

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.068250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:8722155d6e352dd0406091a1f1bb68e8c32a32295465dac5b7b64dc0967e7c23

Observation e34d1d47-aa1f-4d77-8861-8fe651e41ff1 · outbound

This paper cites 2022 , journal=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations 2022 , journal=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.002119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:dfba7ebaf6b53e0d5540dec3b402b65d5f5f94d0834b3c8452329a8b3520d17a

Observation 69f7e846-7b42-45c1-a132-6da1787dae30 · outbound

This paper cites Transactions on Machine Learning Research , year=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Transactions on Machine Learning Research , year=

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.981356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:fc5377dee8ceb9ea1aa07559a4123223327333bbaca294fb1e20f9bfbd9d68a0

Observation 69121aba-0375-44d3-ba70-9283af3100a9 · outbound

This paper cites an unresolved cited work.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-22T04:34:36.991231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:ba7f9578dcef7e83e0f281a94fd62f0312f11206ae0c4ec2c0d36b0703b996b4

Observation 42bdf367-d710-42e8-8428-dcbef79212a7 · outbound

This paper cites Proceedings of the 2022 conference on empirical methods in natural language processing , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2022 conference on empirical methods in natural language processing , pages=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.998239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:0894e0485f83691918444c334874dfb50dbd0dd5140666fc19cf88de82403b85

Observation 5b69e8a9-a4ab-4deb-abb0-dd3659881546 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Advances in Neural Information Processing Systems , volume=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.987886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:a9e58923dd90ba8a8eb4a638cf07f6001ae193413edc55a3494f09f6729ebbe1

Observation a4b800ca-a3da-43b0-a7b1-a1fda654f7fb · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:30:37.615120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:91f17b14699c511b6a7805a970d2cb7c78f79b132048e96a958cd5c5e75c6851

Observation 0184305d-8bb6-4eff-83ea-c80f32d7ecf2 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:59:26.864775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:f4955b2f666092c82217f11eee52d312d8b06fb3aeca02a95c3b3f89a9a19ad1

Observation f26b44fe-ecee-4a67-9ae5-c4390c5e1a6a · outbound

This paper cites Organization of Knowledge and Advanced Technologies.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Organization of Knowledge and Advanced Technologies

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.994809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:cf18a9529ac66597a2720dfe14fc097f1d550ac3c64320bde8912789795eac1e

Observation 8fec285d-6f79-4c4b-af1c-b823e1b66dcf · outbound

This paper cites Efficient LLM Moderation with Multi-Layer Latent Prototypes.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Efficient LLM Moderation with Multi-Layer Latent Prototypes

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-02T04:04:22.908302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:0095e0104c0b2461da9d866e682f91820f14c8720636404c39c2a71a61705c9f

Observation d09d643a-6e56-4864-a8ff-6dafdace5ff8 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.070753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:647256b7564c18d31e0d23e779ddc6261ca7f7aac6880ab8ed10725539c78a4d

Observation c21951ba-222c-4ee3-8d29-9d89bd632016 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.984404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:d40f2350a0dba306fb7f8be0f12865e06314c86e92d5fac09b94123f42e48f9a

Observation 288a5fb8-4d8c-41d7-add8-f2f1fab2587d · outbound

This paper cites arXiv preprint arXiv:2510.06594 , year=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv preprint arXiv:2510.06594 , year=

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.079781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:6595587d90544f62ae94308d0c97a8a2444c68435f345981ca2d61193357d0d9

Observation 1935def0-d11b-4f5c-8a33-228208403b18 · outbound

This paper cites Linearity of Relation Decoding in Transformer Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Linearity of Relation Decoding in Transformer Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.037529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:24d3bb7286e78cf482ed55b475b961c44abb044ac7151697bd325c516cb3dd26

Observation da6ffd10-3a93-4014-835a-f3291df9f26e · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T16:54:37.831311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:b9295060b283957204409b96c4de010f455c6f918e7c5a25474ec484d37f8b1c

Observation 1a266207-ff1e-4b3c-a046-ba8b81cf6fb0 · outbound

This paper cites Companion Proceedings of the ACM on Web Conference 2025 , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Companion Proceedings of the ACM on Web Conference 2025 , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.005418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:6fe4730e6c4be4604c0a1cb6b8fe9e22dbee41c95036a7e205beeed5dbda06ee

Observation 9a5eb6c6-89f2-4c7a-ae93-07f870c29af0 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.941560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:c6593dbfa4efc27655fac78453a6652941fde5220a3a79fcc61ae596070aebd2

Observation 95ba164b-e9b8-4273-83dc-726ed179b39c · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:43:32.339354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:95fcd27f4a4621bd2f4a3c17ad305dfdebfc023f33b683758ebd11897780b47c

Observation fa7e2d14-1c43-4d50-8740-2df308cfcdac · outbound

This paper cites Understanding intermediate layers using linear classifier probes , url =.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Understanding intermediate layers using linear classifier probes , url =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.958245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:c1945dd009cb40bc9c5d3fbc309382fcacba22c977f8f3329e4c230ad4d914c7

Observation af32c80a-6a32-45b6-a975-6977ccb54369 · outbound

This paper cites An introduction to variable and feature selection , url =.

LLM Safety From Within: Detecting Harmful Content with Internal Representations An introduction to variable and feature selection , url =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.938925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:6c180920226c34517e4031cfd7d15063831995551fcf39fef4fea7755c29b22e

Observation ec5f792e-2537-4fb4-a19e-bd90e7521df7 · outbound

This paper cites Qwen3 Technical Report.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Qwen3 Technical Report

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T12:20:23.019486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:2b6f3c608cb5ec404b63601716cab61194ccbbd189075d5435c75fef88d68916

Observation aa71f58b-386e-47f6-9ee3-b2b19877f419 · outbound

This paper cites arXiv e-prints , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv e-prints , pages=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.013438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:85b903a0d9bc68287f062d89bbf5f46a739a2d8bd514bc7e02068ccd77acfb4c

Observation 9b434071-60ab-45dd-94a1-f0891d6c830a · outbound

This paper cites Lightweight Safety Classification Using Pruned Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Lightweight Safety Classification Using Pruned Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.025851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:f3306df553b27cb668f29fa63711840ebc14a45475b2da3d2340dcff6cdce401

Observation 37f1d075-d1a1-44e1-bf1c-d313c220631a · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

LLM Safety From Within: Detecting Harmful Content with Internal Representations ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:17:39.551920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:0722bc618b7c9eb5dd7526c9111987b72d0c108e55c9620e3e988db8302a6ab0

Observation ed317d6a-02e1-4c85-9ecf-89b825591580 · outbound

This paper cites Generative or Discriminative? Revisiting Text Classification in the Era of Transformers.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Generative or Discriminative? Revisiting Text Classification in the Era of Transformers

Reference 47

Resolution
verified exact
doi, observed 2026-05-10T04:35:16.476296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:276e3810ac31af4ccc8c2b9c31f913b9daadc9d4fe513aa0056752178239c083

Observation 4dd6b657-03f0-4045-9038-4df151e4383b · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations The Thirteenth International Conference on Learning Representations , year=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.023071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:0b44bd0bf97287b2d3150541114cd078824cf4d52620b67864f862c90f067aff

Observation 9e0f907f-bdd2-42f2-8de5-68a3b7994bf9 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.948506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:ec9bc898e3f921a7ef8a60709de71c9520d444e2e66a4d945f1ff81149a4a9f8

Observation 134b24df-8dfe-496c-a19f-abc97a7ff2f8 · outbound

This paper cites IEEE Robotics and Automation Letters , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations IEEE Robotics and Automation Letters , volume=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.978073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:f9b70d8f1864b2ecc896399a23ba55430b70da28941f6276bd816dffd6c77b4b

Observation dff3646d-ff43-490f-af8f-75e286bb77b7 · outbound

This paper cites RDI: An adversarial robustness evaluation metric for deep neural networks based on model statistical features.

LLM Safety From Within: Detecting Harmful Content with Internal Representations RDI: An adversarial robustness evaluation metric for deep neural networks based on model statistical features

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.060041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:0c08e64d6beab53cef52b87c01e065bc7491ba3a266e97d5c135bb9495d7e223

Observation 0352f629-72d9-4450-8d61-fe393b41a693 · outbound

This paper cites Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.944899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:4724e7b29a1dcb59cd4f90cca46d2b27728f8af1cf7b74e981b1c543cf531399

Observation 25396c11-0073-4387-9657-1dcad9e06309 · outbound

This paper cites Scaling laws for neural language models , url =.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Scaling laws for neural language models , url =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.031134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:d22f3fcfa66173d3fdc607be800e79b4af8706434b73143f17cf78d5edf4c7ea

Observation ecedd5a7-16f4-4a84-b195-84538c5d991d · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T02:34:44.407176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:996605ed69ead4eb2d7cc574b3ea39fa15b2abfd8c68090d94b1e690e4cae7ad

Observation 4b56e116-1bfb-49f2-85bf-d149403d1cdf · outbound

This paper cites author=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations author=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.027719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:777f2ddd84925fa67f9c1996eb8c201b1cbf0655d5c4f188325b5ba28bd3ea49

Observation f437287a-2719-4f3a-9179-e74fe41ecda4 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T12:20:23.010454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:84917fec681cc6921076c4d3a47059c174a68884de39f5ec6ef99394630ae71e

Observation 0c8b0a62-c9a2-44f9-b6b3-d4a9fbeb5989 · outbound

This paper cites 2025 , howpublished =.

LLM Safety From Within: Detecting Harmful Content with Internal Representations 2025 , howpublished =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.302967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:081f46b1a97169b1319a1d33cb75da1f57fe84971d2354d16cdb400282e526ba

Observation 9e552a06-6d22-4761-a2d0-2b53260aa24c · outbound

This paper cites 2025 , howpublished =.

LLM Safety From Within: Detecting Harmful Content with Internal Representations 2025 , howpublished =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:36.951802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:0325d2abc9a274c6c483a4cf53c9c20c8a405d93ce10dc7293293b4199ab2d2f

Observation 7b9dbb1e-d773-4a52-ac22-4b31eac9f159 · outbound

This paper cites an unresolved cited work.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Unresolved cited work

Reference 59

Resolution
parse uncertain
raw_fallback, observed 2026-05-22T04:34:36.955330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:f2a4f32ba67dac60025b3404983f745dbbbe4b8984295158d160a7278ae8aee9

Observation 4ffef57a-39c4-491a-8344-0b79e4db3752 · outbound

This paper cites Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.048764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:1319f1eeda0a1cd53645c273923fd4dff70facc8d57e24e1beec0c4648d52e8f

Observation 7426ea55-99ba-44c8-a9e5-397ca6e2f056 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

LLM Safety From Within: Detecting Harmful Content with Internal Representations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T12:20:23.040333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:8e7812e89ea244dd546e9ab0a6955b5e8253658aad9f4b3f7364b5196e624a63

Observation 96928c83-816f-4f73-adf8-bb0302d28415 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

LLM Safety From Within: Detecting Harmful Content with Internal Representations ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:08:10.048088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:d383c66df2a2e1ec49b9a35dc2a980764cc44852f78b1e006e1ca408b0cc44d2

Observation 190f9fcb-274b-4fad-91b2-00244340b913 · outbound

This paper cites arXiv preprint arXiv:2508.03550 , year=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv preprint arXiv:2508.03550 , year=

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.045949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:776bd8547f304f56512cb5fbd19837610dc5c1dcea2077b4955a4a35b16d0dec

Observation c401dee2-40a4-4c16-a905-26cd98ef813d · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.294906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:bbeb8f40f4674fc868ac7aaf40b6f83d9e6009d83524dc0cd7a171e05070449d

Observation fd9f2d22-03f3-4eac-a64c-2d087444a635 · outbound

This paper cites BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding.

LLM Safety From Within: Detecting Harmful Content with Internal Representations BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 65

Resolution
verified exact
doi, observed 2026-05-10T04:35:16.486832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:a7c8bebb67c7e2e7a5f64c9e1c736325f02659765836a04955fab999b9e43ad6

Observation 5b760b1f-10fb-4087-a3cf-f513f2d6604f · outbound

This paper cites 2019 , eprint=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations 2019 , eprint=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.298967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:fca22dda9ec4b1182bc0828979931a84dc0913dbdda25d5af4a3b19a0d906fd9

Observation d0c1504c-2ded-4914-9a9f-f6a8b85423c4 · outbound

This paper cites Hate speech detection and racial bias mitigation in social media based on bert model.PLOS ONE, 15(8):1–26, 08 2020.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Hate speech detection and racial bias mitigation in social media based on bert model.PLOS ONE, 15(8):1–26, 08 2020

Reference 67

Resolution
metadata mismatch
doi, observed 2026-05-10T04:35:16.482488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:1d70dfadbef6a1a032d1b64e067e3a93046b89110a441aba49b842d34a177d23

Observation cc4b602a-2457-4275-b6ea-74d178d22fdf · outbound

This paper cites 2021 , isbn =.

LLM Safety From Within: Detecting Harmful Content with Internal Representations 2021 , isbn =

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:35:16.480828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:87453c4cca8fad8e709374b7cb9ef7f0cefa7d3185213313dcec9daa57f4a05a

Observation 10d91182-2d94-4175-8282-2944ecafb995 · outbound

This paper cites H ate BERT : Retraining BERT for Abusive Language Detection in E nglish.

LLM Safety From Within: Detecting Harmful Content with Internal Representations H ate BERT : Retraining BERT for Abusive Language Detection in E nglish

Reference 69

Resolution
verified exact
doi, observed 2026-05-10T04:35:16.478300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:a751fd7f7a60c2dbc1fbc2b211361ab65c90c847f469f2779808d92b73074c97

Observation 02b07d0d-af88-48fd-8079-717a37ffa34a · outbound

This paper cites 2024 , eprint=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations 2024 , eprint=

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.306299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:1f34e8d97aff7be0ee69c6e6833864bf1412a29968443d887b96d9a6a9531e91

Observation 06a6ec07-824a-4b46-8aa8-5c30d572501c · outbound

This paper cites and Tay, Yi and Sorensen, Jeffrey and Gupta, Jai and Metzler, Donald and Vasserman, Lucy , title =.

LLM Safety From Within: Detecting Harmful Content with Internal Representations and Tay, Yi and Sorensen, Jeffrey and Gupta, Jai and Metzler, Donald and Vasserman, Lucy , title =

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T04:35:16.485012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:81331c7009f8e2849fa1df35d4b0a4d0c9f1de7213c6d697e99f4a98a44572e4

Observation 5abbd1ea-32e1-4861-b346-9b5e7c58d9c5 · outbound

This paper cites 2025 , eprint=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations 2025 , eprint=

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:34:37.034244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:b6c8eba2a1bc3e7ad9f3ae5cb9cec8857bd6fe3075375522a02064293168b1b5

Observation 0421a101-051c-442a-a6b4-4b3a122c85d1 · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.309711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:18639553879646846c60ec8b7b125129b89f693794bb04f1805ea78e8cb55235

Observation 1e9acd6d-91b0-4870-a830-94de92d2b6cd · outbound

This paper cites PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages.

LLM Safety From Within: Detecting Harmful Content with Internal Representations PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.043286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:d8bb4537a58b137bb38cd03f2ff8c58e9bb51f4fa983f083746db073309c08bb

Observation 1a853100-0621-4b91-8431-916beeb54807 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Linear Representations of Sentiment in Large Language Models

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:43:06.705954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:7ef47e1a3ca311693604504c742d182871681fa464f6761dcb9c1ecddda1fac6

Observation 36b3f119-44c3-4de0-8410-ade08be64a1a · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

LLM Safety From Within: Detecting Harmful Content with Internal Representations The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T19:35:39.400821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:48948d591b93797ad99e08171b39f1e24bd481f89d72f25eca5812b145f0de3a

Observation 279548d1-4dae-4d91-9093-32759b537cbf · outbound

This paper cites Proceedings of the 28th ACM international conference on information and knowledge management , pages=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Proceedings of the 28th ACM international conference on information and knowledge management , pages=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.315979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:c5e3caffd6c5cb4e41b637f18d80e5a78dc04a3ae5e29d5e174dfca8f468b550

Observation f2aa0f24-f7bf-4c7f-9eb9-2c042ce9d75b · outbound

This paper cites arXiv preprint arXiv:2510.18081 , year=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations arXiv preprint arXiv:2510.18081 , year=

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:20:23.065573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:1f9eb4474aca47eda3da20370118091044ff8441a6102f6596ebbbeaf9982859

Observation 9b0a7544-a00d-4285-9474-24c935d534b5 · outbound

This paper cites Curvalid: Geometrically-guided adversarial prompt detection.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Curvalid: Geometrically-guided adversarial prompt detection

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.073944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:c804b1d612585250cd32338f01b641e014a8c55dfe886af78aa612a8215b0adc

Observation 79fa70d3-2184-45cb-b67e-8ed50c7fd44f · outbound

This paper cites Neural networks , volume=.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Neural networks , volume=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T04:36:04.312630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:4065ed0888588d271247d5d7ef9b1bbce7e603942d57855747dfdbcab6597dc7

Pith citing papers

Observation 85e0dbd5-f316-43eb-a156-6af8652f9f7f · inbound

MINER: Mining Multimodal Internal Representation for Efficient Retrieval cites this paper.

MINER: Mining Multimodal Internal Representation for Efficient Retrieval LLM Safety From Within: Detecting Harmful Content with Internal Representations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:06:10.607800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T12:40:33.437364Z digest=sha256:3552964322e10c6101209e473d8041a9f1a8289f61ec24c8b3281bb37343addd

Observation 89bd88d0-aa8e-4cca-80eb-23bae07202bb · inbound

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems cites this paper.

LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems LLM Safety From Within: Detecting Harmful Content with Internal Representations

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:01:05.907482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T04:58:37.780195Z digest=sha256:0e074b025e349f309fea12646cb2aa7f2e4b5b3282eafae5da54026a48180b78

Observation 07051512-4efa-4e18-b1c7-73f64e217c15 · inbound

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue cites this paper.

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue LLM Safety From Within: Detecting Harmful Content with Internal Representations

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.762580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:28:25.577739Z digest=sha256:f1c696874397d00960eebfdf3c7ccdbd309c331a76da3e9c1c16e2746e5d52bc