Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:16.539220Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 43 inbound Pith citation observations for arXiv:2505.23836.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:16.539220Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:19:51.357826Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
13 of 13 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation ee4c52f0-ab01-483b-a8ab-6f74e0b8d19d · outbound
Large Language Models Often Know When They Are Being Evaluated Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c42e766-0c9e-4716-8ed9-04bf6b419834 · outbound
Large Language Models Often Know When They Are Being Evaluated Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8d2f4893-47b3-4137-af3c-f0aa5f4dec6e · outbound
Large Language Models Often Know When They Are Being Evaluated SafetyBench: Evaluating the Safety of Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd404eaa-09a4-4f82-bc8e-33d8fa5fc2ec · outbound
Large Language Models Often Know When They Are Being Evaluated Fictional Scenario
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 62c685d9-cf3e-49f8-872e-58ae93b95136 · outbound
Large Language Models Often Know When They Are Being Evaluated test" or
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bd4373ad-587b-498f-bdce-17f0110f9107 · outbound
Large Language Models Often Know When They Are Being Evaluated Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ef4e8172-1c6c-4446-bae3-c66de5f65b6e · outbound
Large Language Models Often Know When They Are Being Evaluated Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 73cfae9e-ee65-4db2-a7f4-e9aa548beb77 · outbound
Large Language Models Often Know When They Are Being Evaluated Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9d3697b5-9296-4033-acaa-e7c8039ac2e7 · outbound
Large Language Models Often Know When They Are Being Evaluated Figure 18: The UI used by the authors to annotate the transcripts and create the human baseline
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3d05b928-c056-401c-bd28-2c7c2c051e3f · outbound
Large Language Models Often Know When They Are Being Evaluated ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d8ff20-356b-4be2-9aba-bb9b38955dea · outbound
Large Language Models Often Know When They Are Being Evaluated ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcf12358-b2e3-48bf-86b9-bff50897266d · outbound
Large Language Models Often Know When They Are Being Evaluated XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5881551b-1fe8-446c-b433-fd14e10458b8 · outbound
Large Language Models Often Know When They Are Being Evaluated AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71751645-d83b-4346-ab32-026e5a3c704d · inbound
AI Awareness Large Language Models Often Know When They Are Being Evaluated
Reference 172
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a33385-10e3-4e59-8446-4f0ae2d562a3 · inbound
The California Report on Frontier AI Policy Large Language Models Often Know When They Are Being Evaluated
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32a4d493-fc2c-4b5f-bca2-5444675ec82b · inbound
Subversion via Focal Points: Investigating Collusion in LLM Monitoring Large Language Models Often Know When They Are Being Evaluated
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 014e48a2-b814-4089-aba3-d37313ff299e · inbound
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language Large Language Models Often Know When They Are Being Evaluated
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d196b84c-0b27-4bf2-a575-2e893d947be3 · inbound
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Large Language Models Often Know When They Are Being Evaluated
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a276a77f-28fc-4934-81cf-438322bf642c · inbound
Are LLM Belief Updates Consistent with Bayes' Theorem? Large Language Models Often Know When They Are Being Evaluated
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca63ec3a-d0b1-4d4c-a9af-07197cc22118 · inbound
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases Large Language Models Often Know When They Are Being Evaluated
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6fff0f1-dedf-4da2-bbe2-afcc30102295 · inbound
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Large Language Models Often Know When They Are Being Evaluated
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fc8eaeac-ef49-45ef-ad60-8fdfcbb3a357 · inbound
An Independent Safety Evaluation of Kimi K2.5 Large Language Models Often Know When They Are Being Evaluated
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0a7a6c42-8f3b-4ca0-8a9d-a7f391032f7d · inbound
Simulating the Evolution of Alignment and Values in Machine Intelligence Large Language Models Often Know When They Are Being Evaluated
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e931c58f-1d37-4a12-8957-f3dacf5c1eee · inbound
Honeypot Protocol Large Language Models Often Know When They Are Being Evaluated
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a5726432-a35f-4568-9ac5-ff01aa23a448 · inbound
Risk Reporting for Developers' Internal AI Model Use Large Language Models Often Know When They Are Being Evaluated
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 11820b76-29ca-487c-bbae-8368eaea64ef · inbound
Towards Understanding Specification Gaming in Reasoning Models Large Language Models Often Know When They Are Being Evaluated
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3ffaa199-acf2-41bc-b72f-3b1994ed3d86 · inbound
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use Large Language Models Often Know When They Are Being Evaluated
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ec45d5b5-4ded-4809-9c16-e34ed3515f9b · inbound
Evaluation Awareness in Language Models Has Limited Effect on Behaviour Large Language Models Often Know When They Are Being Evaluated
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b4e8bf0d-15fc-494a-9c0d-aee4dc2e0600 · inbound
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors Large Language Models Often Know When They Are Being Evaluated
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2c504cd9-7b1a-42f8-a979-2cb3ec681887 · inbound
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels Large Language Models Often Know When They Are Being Evaluated
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 29f87113-88d6-47b3-a880-7cfefb06b50d · inbound
The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime Large Language Models Often Know When They Are Being Evaluated
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d12d4985-7841-4d48-aacc-95684c94c188 · inbound
Naturalistic measure of social norms alignment Large Language Models Often Know When They Are Being Evaluated
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4f8a1ee2-2cfa-4ed0-b5cb-48c5b5f25f0c · inbound
Consistency Training while Mitigating Obfuscation via Rate Matching Large Language Models Often Know When They Are Being Evaluated
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation adfc2b2e-db10-4ab5-be71-542087ee551f · inbound
Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents Large Language Models Often Know When They Are Being Evaluated
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6a8b9bfd-f0d7-44cf-8261-44acad1fd5d7 · inbound
LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Large Language Models Often Know When They Are Being Evaluated
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5eb1a384-aa8d-4722-9772-0ed7bf494874 · inbound
Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Large Language Models Often Know When They Are Being Evaluated
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 62bfab90-7019-4ae2-ab7c-41c8a09278d0 · inbound
Sycophancy Towards Researchers Drives Performative Misalignment Large Language Models Often Know When They Are Being Evaluated
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 09e0708c-53d6-47e4-906c-e6fc347d0fe3 · inbound
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs Large Language Models Often Know When They Are Being Evaluated
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 37fd511b-e396-443e-ae66-40f8d5f092be · inbound
Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization Large Language Models Often Know When They Are Being Evaluated
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 24a2c7fb-6793-4a10-95aa-d4e0f3d12700 · inbound
Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs Large Language Models Often Know When They Are Being Evaluated
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 512610eb-fc91-4a50-b0f6-ce0079167855 · inbound
Evaluation Awareness Is Not One Capability: Evidence from Open Language Models Large Language Models Often Know When They Are Being Evaluated
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3242f6d5-421a-45a4-9aad-0d419c6ac36e · inbound
Defeat Devices in AI Systems Large Language Models Often Know When They Are Being Evaluated
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 71ad9311-cc8c-4b1f-a7cc-ac1b262d95d4 · inbound
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Large Language Models Often Know When They Are Being Evaluated
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a59c605b-f5b8-4386-9a4f-7586d3f74aab · inbound
Predicting LLM Safety Before Release by Simulating Deployment Large Language Models Often Know When They Are Being Evaluated
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3f181c66-37a7-4bc2-b6f4-0822b3871b4b · inbound
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Large Language Models Often Know When They Are Being Evaluated
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0632d1-31cc-4481-97af-30d0b63d7694 · inbound
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models Large Language Models Often Know When They Are Being Evaluated
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73131b97-d4f1-457a-844c-2339afd33228 · inbound
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory Large Language Models Often Know When They Are Being Evaluated
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0cd2c7f-7416-43b1-b555-412c771768e6 · inbound
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Large Language Models Often Know When They Are Being Evaluated
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 592725c1-a39b-45e3-b101-7b47c9da458d · inbound
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models Large Language Models Often Know When They Are Being Evaluated
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e48d91-61f2-4ac3-9bcb-31e893026e59 · inbound
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? Large Language Models Often Know When They Are Being Evaluated
Reference 168
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61b0ff3f-f4ab-4b02-87d9-3ba13e643245 · inbound
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Large Language Models Often Know When They Are Being Evaluated
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd384b3-2bfc-4a56-898a-426a4397d627 · inbound
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation Large Language Models Often Know When They Are Being Evaluated
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b4dd25-ddf0-4615-be23-e88bd3efb9f5 · inbound
Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance Large Language Models Often Know When They Are Being Evaluated
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e85087d-387f-4c75-89db-e0fb1db3f69d · inbound
Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance Large Language Models Often Know When They Are Being Evaluated
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 141c6d5b-da9a-4c0f-96ff-bb6c0d23d9a0 · inbound
Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation Large Language Models Often Know When They Are Being Evaluated
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 380638a4-fddc-4bc8-858d-53a9ee46c614 · inbound
A Probe Direction Is a Property of Its Prompt Large Language Models Often Know When They Are Being Evaluated
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.