Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T21:06:56.135403Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 1 inbound Pith citation observation for arXiv:2502.05242.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T21:06:56.135403Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:01:59.215909Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T23:02:04.281602Z
91 of 91 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f3e8f967-b44d-4ec1-bfce-2738e74c1b1f · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Generative AI Text Classification using Ensemble LLM Approaches
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e63ebdc5-3aad-4860-8b8c-4fcc8b8591fb · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0e013a8-1d8a-4685-8774-759eba133331 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Internal State of an LLM Knows When It's Lying
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457ddc55-500e-4a0c-be87-d9b5acc5e94d · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Transparency and explainability of ai systems: ethical guidelines in practice
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b937a963-1c1c-4573-bf94-74451f211016 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Spectrally-normalized margin bounds for neural networks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd51a220-ef09-4d7a-aed8-ca8c6dcd3fca · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Language models can explain neurons in language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff633a2-2ed4-4a3e-9fb4-c940d9671e89 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring High-Dimension Human Value Representation in Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f4e9428-e1ba-410b-b894-69a5806bac97 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Internlm2 technical report, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4f9efca-d4c3-4300-863c-4cf5ea4377c2 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Redunet: A white-box deep network from the principle of maximizing rate reduction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae97825f-34b0-4041-aa53-5e11bac06f2d · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring SelfIE: Self-Interpretation of Large Language Model Embeddings
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65862f5e-6e05-4c89-8962-8c35adb1434c · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring A simple framework for contrastive learning of visual representations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d84af8-53d5-4908-8fbe-29a7d337e092 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Reasoning Models Don't Always Say What They Think
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3eb7f9c-41eb-4d33-8e1b-1c3bb0803fbc · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Token Prediction as Implicit Classification to Identify LLM-Generated Text
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5939e0aa-5dd2-4819-87c4-6cee38a1f6ef · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Measuring generalization with optimal transport
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 501c34ff-aac9-4645-b297-50d7bd509842 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Training Verifiers to Solve Math Word Problems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f210b8b-a74b-4846-8cc4-0969bdab1572 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Opencompass: A universal evaluation platform for foundation models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a558fa4f-2a39-4149-80df-bd468a5a5042 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a712fe56-9fc0-4031-9daf-03316b48d32b · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e02df84-da1c-4178-a019-39ed782397c8 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80e24131-ca75-4d12-947c-9c2e86df1e54 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Scaling and evaluating sparse autoencoders
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6e1392-4cdc-4298-9c06-0f4bd94eaef8 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0187d06a-8394-441e-9bb2-d9e927e54d26 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Alignment faking in large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b469aef6-7e91-495f-ad40-23df6a2e9795 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Dimensionality reduction by learning an invariant mapping
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0fa4ebb-603c-496e-b878-aafe3fdfd14b · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Masked autoencoders are scalable vision learners
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bcd3c8-ce88-43ee-8aff-f1ca7ead4c4c · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Momentum contrast for unsupervised visual representation learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688d65ed-4bff-4110-a7af-35d28a44760d · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Measuring Massive Multitask Language Understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b175704b-317e-4612-ad96-c1a7e576cdfc · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Measuring mathematical problem solving with the math dataset
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7ac45f8-93e2-4db4-8b6e-513633c2f52a · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring LoRA: Low-Rank Adaptation of Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ff69cd-e2d8-495a-bf29-c5ad0093834c · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173d7867-b738-4c0a-9787-e087706fc231 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Can Large Language Models Explain Themselves? A Study of LLM-Generated Self-Explanations
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10957a03-58f7-494b-875c-b555402d289d · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35dd0db-4a15-410d-8ace-4a2453be7f10 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Anthropic: Responsible scaling policy
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e83a56d-4123-4e6b-b529-66b410741f27 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Platonic Representation Hypothesis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83447c0a-c73c-4565-8128-9ab08938f379 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Klanderman, and William J Rucklidge
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4e41a44-f263-41ee-84a9-6447e16b6d3b · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70aa9370-bc40-4292-9e1d-49f62f31b051 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adcf540f-b082-4a3e-8918-95a7bd6c7d59 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Mistral 7B
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d035885d-8c60-4038-ad07-fdf4cbde629f · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13866d13-644a-4719-b597-dc36bef6e648 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Explainable artifi- cial intelligence for mental health through transparency and interpretability for understandability
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae57fff5-306c-420d-9be1-75a143f4fa31 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Theoretical Analysis of Weak-to-Strong Generalization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e56eae66-2a65-4723-84a4-b2a9ea64b870 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff0e8ae-7aed-43ef-b4ef-f79cf43fff2e · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Inference- time intervention: Eliciting truthful answers from a language model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3b5c0a1c-1092-451e-abaf-5fcbafeda044 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f68183-74c7-4e22-bd63-d35465225cf9 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f0ad26-1ff9-4a3a-8c07-5e261d86327f · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce0c3e3-8ff9-4f56-949e-5081cf4e12d3 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Large language models in finance: A survey
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa921c34-69a1-4f13-bb2a-16033b96982f · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fe70bf5-43ea-4c1f-9120-0d1084c5f9b8 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Latent guard: a safety framework for text-to-image generation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a83b41f-745d-483b-9005-e7157117c742 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 036de2d0-4613-491c-bb9f-b3cc5f5c938a · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Is the MMI Criterion Necessary for Interpretability? Degenerating Non-causal Features to Plain Noise for Self-Rationalization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a900ae02-3466-4912-8423-ea555e038e6d · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring MGR: Multi-generator Based Rationalization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e06cb80-918d-4863-8234-b1daffbddaf5 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring D-separation for causal self-explanation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation df5fed00-963e-4362-b243-b73676f467f7 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Efficient detection of toxic prompts in large language models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16164951-f8c2-42c2-95a9-11c2a14fd8e2 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Are self-explanations from large language models faithful? In Findings of the Association for Computational Linguistics ACL 2024, pages 295–337, 2024
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 446c31d7-65c7-4a8a-b4c3-89caa757e468 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Introducing llama 3.1: Our most capable models to date
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d218019-ad92-4d0e-9b10-81277d858604 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Large language models in healthcare and medical domain: A review
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20430a8d-5418-4d9d-9714-48fe056b05aa · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring interpreting gpt: the logit lens, 2020
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 225d2322-a4f4-4637-8b8b-803d0835fc56 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Show Your Work: Scratchpads for Intermediate Computation with Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2653e440-a101-4d97-a749-9c18c20fe79a · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Representation Learning with Contrastive Predictive Coding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77650200-95ad-4254-ae66-b393f8e2d986 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f79b94-4509-4a6c-9844-3692db5eff03 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Training language models to follow instructions with human feedback
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df7f434-40a5-467f-824c-394f877165eb · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747cf1a8-5e93-4708-8a45-16bfebbf9b3c · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Learning transferable visual models from natural language supervision
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0bdbcb-74c2-45ee-80d8-7e72561f0d4f · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Representation Noising: A Defence Mechanism Against Harmful Finetuning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1aa9788-3f59-43b1-9fcc-dc77ff0d4fc9 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0568c9d-14e2-486f-b012-9c06393075b6 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The effective rank: A measure of effective dimensionality
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60750dfb-be55-4241-815e-7f1a073573f9 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring SocialIQA: Commonsense Reasoning about Social Interactions
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b947ce92-bb8f-4fc5-bbe7-cbb7fd9dc66f · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Facenet: A unified embedding for face recognition and clustering
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17821f04-8ab1-4875-9197-6ea97b00a04f · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Lmfingerprints: Visual explanations of language model embedding spaces through layerwise contextualization scores
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2c8f106-6e1b-4c54-961a-ecbcd71885f5 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring k-variance: A clustered notion of variance
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1de0eab6-ba41-4b39-ab5e-adb1a6deb80d · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Gemma 2: Improving Open Language Models at a Practical Size
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e19d31c-1e90-4440-ad87-d31159744ba7 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Daniel Freeman, Theodore R
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d11c2a-7849-4444-b3f6-cb06558089a2 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd3ee770-4322-4bd8-9d30-f2c3e99d076c · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Optimal transport: old and new, volume 338
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5f683d-ee24-4e37-bb56-4e66c8530de7 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring The wasserstein distances
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8d9cf8e-f1cc-4754-95ea-77d0287e45b6 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30e8831-9260-4a73-817d-aa3178454530 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Chain-of-thought prompting elicits reasoning in large language models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1bae7d8-6fc3-4d92-8aa1-a01fabec5b31 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Diff-erank: A novel rank-based metric for evaluating large language models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98241810-e6fc-4da6-b8c2-fdd71696400b · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring ReFT: Representation Finetuning for Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c478ef5-0b34-4c83-a5d5-0e163c9ce198 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Good Idea or Not, Representation of LLM Could Tell
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4abf341-265d-49b5-9dd6-72f6a8734d07 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Qwen2.5 Technical Report
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f6b988-4f04-4cc0-a4eb-880f270b3154 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring A fingerprint for large language models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091dd076-9a66-4127-95df-8f08f6cecf83 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring LoFiT: Localized Fine-tuning on LLM Representations
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab498e3-2f81-4167-bdea-0c80fc1fe5dc · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Barlow twins: Self- supervised learning via redundancy reduction
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d647cc87-41ae-40f1-8821-278ab1fcf4b6 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Similar Data Points Identification with LLM: A Human-in-the-loop Strategy Using Summarization and Hidden State Insights
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcbaeea0-a2a4-4a7a-b6d0-5364c23af323 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring REAL: Response Embedding-based Alignment for LLMs
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62028e91-3ff0-4b67-8729-f27302ce3255 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring REEF: Representation Encoding Fingerprints for Large Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d8497d-01d9-4a05-9ebb-a9209b335a59 · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a13672-7dec-4bf3-a6fa-701fb7cf5a0d · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Improving alignment and robustness with circuit breakers
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e295cd8-8b93-4bf9-866d-ec29bbb9cc8e · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring [ 62] disentangles LLMs’ awareness of fairness and privacy by deactivating the entangled neurons in representations
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4ef902d-32e2-4788-bb28-1158b52b904a · outbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring safe" or
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0bdfb9f2-0a78-4996-bd33-6a36b60098b2 · inbound
Position: Intelligent Coding Systems Should Write Programs with Justifications Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.