Pith. sign in

Paper Citation Record · LEDGER

Large Language Models Often Know When They Are Being Evaluated

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 43 inbound Pith citation observations for arXiv:2505.23836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23836 v3

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:16.539220Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:19:51.357826Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved7
  • parse uncertain2
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ee4c52f0-ab01-483b-a8ab-6f74e0b8d19d · outbound

This paper cites an unresolved cited work.

Large Language Models Often Know When They Are Being Evaluated Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:18.050051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:15.989984Z digest=sha256:4c561efc1e0e59c67f95a898b7de004bdf84d6c1cf671f67912a183bd4491971

Observation 4c42e766-0c9e-4716-8ed9-04bf6b419834 · outbound

This paper cites an unresolved cited work.

Large Language Models Often Know When They Are Being Evaluated Unresolved cited work

Reference 2

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T13:16:17.865736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:16.082916Z digest=sha256:d77a690dfa25d42e5d6479b453f3b92d2ce7c26f8a90bb6a0d2753fd0142660d

Observation 8d2f4893-47b3-4137-af3c-f0aa5f4dec6e · outbound

This paper cites SafetyBench: Evaluating the Safety of Large Language Models.

Large Language Models Often Know When They Are Being Evaluated SafetyBench: Evaluating the Safety of Large Language Models

Reference 3

Resolution
malformed identifier
no resolver link, observed 2026-08-07T13:16:15.735454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.735454Z digest=sha256:2f424cdeca04c40421d785c9c7899c2422c6cd47f330963a4a66e218603d39d5

Observation bd404eaa-09a4-4f82-bc8e-33d8fa5fc2ec · outbound

This paper cites Fictional Scenario.

Large Language Models Often Know When They Are Being Evaluated Fictional Scenario

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:17.460904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:16.273381Z digest=sha256:57a9404ebb03141d28271edbb12d7363a3574c2f9ca4d5edf7c6b18683be9594

Observation 62c685d9-cf3e-49f8-872e-58ae93b95136 · outbound

This paper cites test" or.

Large Language Models Often Know When They Are Being Evaluated test" or

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:18.288641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:15.896241Z digest=sha256:03140ce1addba03a6faaa50f222f4850da7f1605ed69d737056d8ebc0cc74c14

Observation bd4373ad-587b-498f-bdce-17f0110f9107 · outbound

This paper cites an unresolved cited work.

Large Language Models Often Know When They Are Being Evaluated Unresolved cited work

Reference 9

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T13:16:17.641327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:16.176113Z digest=sha256:b3a45ef0d17d75bbec5da3e8a79182e488e89be1c0326216669527e6ae12619b

Observation ef4e8172-1c6c-4446-bae3-c66de5f65b6e · outbound

This paper cites an unresolved cited work.

Large Language Models Often Know When They Are Being Evaluated Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.303560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:16.368296Z digest=sha256:20b0d6bfcb747820aa8d62de0fcc0a441f4faa98efc692a1ad99028bb2388367

Observation 73cfae9e-ee65-4db2-a7f4-e9aa548beb77 · outbound

This paper cites an unresolved cited work.

Large Language Models Often Know When They Are Being Evaluated Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:16:17.134569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:16.465156Z digest=sha256:f7ef2e4a41dfb7b65130bfe06894db4b1e718ee494178f68c717b185dd13a9b7

Observation 9d3697b5-9296-4033-acaa-e7c8039ac2e7 · outbound

This paper cites Figure 18: The UI used by the authors to annotate the transcripts and create the human baseline.

Large Language Models Often Know When They Are Being Evaluated Figure 18: The UI used by the authors to annotate the transcripts and create the human baseline

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:16.971288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T13:16:16.539220Z digest=sha256:5a3a7530ca6b065dd2e99f1000bb38dac2e6e12420bb5934e71777c3dc872737

Observation 3d05b928-c056-401c-bd28-2c7c2c051e3f · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Large Language Models Often Know When They Are Being Evaluated ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.857298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.857298Z digest=sha256:ab08e700c3ac7c9fcc5aefca3456b5a991bb69710adb932555881492a742b374

Observation 96d8ff20-356b-4be2-9aba-bb9b38955dea · outbound

This paper cites ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation.

Large Language Models Often Know When They Are Being Evaluated ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.795060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.795060Z digest=sha256:39f0590b61c03b7d889251ea67dd21ef99ad3e1f0d4a7ce56592c208e014bc4f

Observation dcf12358-b2e3-48bf-86b9-bff50897266d · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Large Language Models Often Know When They Are Being Evaluated XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.583967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.583967Z digest=sha256:1b00700596d504e3b04a7308a1b97949f205a75915880d5fd1dc45fdb300d258

Observation 5881551b-1fe8-446c-b433-fd14e10458b8 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Large Language Models Often Know When They Are Being Evaluated AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.669721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.669721Z digest=sha256:7caadfcfe38f3545d7cda8b1143b03949ce6a055110ed8166c002825c38b2dc9

Pith citing papers

Observation 71751645-d83b-4346-ab32-026e5a3c704d · inbound

AI Awareness cites this paper.

AI Awareness Large Language Models Often Know When They Are Being Evaluated

Reference 172

Resolution
unresolved
no resolver link, observed 2026-08-16T10:19:51.357826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:19:51.357826Z digest=sha256:33bdf661acebf3cd17b133338c6e7004618ff7b776279ddd29fd7b6bf206d138

Observation e5a33385-10e3-4e59-8446-4f0ae2d562a3 · inbound

The California Report on Frontier AI Policy cites this paper.

The California Report on Frontier AI Policy Large Language Models Often Know When They Are Being Evaluated

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:40.968945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:40.968945Z digest=sha256:7bcab5fd3b4220348ccd83147e9a85be05fe42fa5aa66d31b6e2d94d99982d3e

Observation 32a4d493-fc2c-4b5f-bca2-5444675ec82b · inbound

Subversion via Focal Points: Investigating Collusion in LLM Monitoring cites this paper.

Subversion via Focal Points: Investigating Collusion in LLM Monitoring Large Language Models Often Know When They Are Being Evaluated

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:52:29.855907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:52:29.855907Z digest=sha256:ccec876ea5e901f724a28fb6ce72b16ffa3c36ab4c18616158bb9585b652e4a0

Observation 014e48a2-b814-4089-aba3-d37313ff299e · inbound

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language cites this paper.

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language Large Language Models Often Know When They Are Being Evaluated

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:18:15.522929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:18:15.522929Z digest=sha256:b48d37eb97b84895e6bf776676f7527ca092928d9f342347a6f824035bcdfa3a

Observation d196b84c-0b27-4bf2-a575-2e893d947be3 · inbound

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework cites this paper.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Large Language Models Often Know When They Are Being Evaluated

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.667793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.667793Z digest=sha256:2e33bf1e8507b18d02444450ee5ff07fefecb25b5145a0ba0b371702d9d2e1a1

Observation a276a77f-28fc-4934-81cf-438322bf642c · inbound

Are LLM Belief Updates Consistent with Bayes' Theorem? cites this paper.

Are LLM Belief Updates Consistent with Bayes' Theorem? Large Language Models Often Know When They Are Being Evaluated

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:24:06.722915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:24:06.722915Z digest=sha256:2126110e0be984129ddd777ee108ef692207ce4c2f9b609bd4f491b5ce2a280a

Observation ca63ec3a-d0b1-4d4c-a9af-07197cc22118 · inbound

Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases cites this paper.

Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases Large Language Models Often Know When They Are Being Evaluated

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:43.213404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:16:43.213404Z digest=sha256:320114072be6a32cef2700549094f563ee034e31eeb9b073a4f09c5579c860af

Observation d6fff0f1-dedf-4da2-bbe2-afcc30102295 · inbound

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies cites this paper.

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Large Language Models Often Know When They Are Being Evaluated

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:31:29.441230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T14:31:11.974395Z digest=sha256:88bb75e507d4c616761c7b8c1f051a03fb199ddf1de8fccd3ac5de55dde497dc

Observation fc8eaeac-ef49-45ef-ad60-8fdfcbb3a357 · inbound

An Independent Safety Evaluation of Kimi K2.5 cites this paper.

An Independent Safety Evaluation of Kimi K2.5 Large Language Models Often Know When They Are Being Evaluated

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:43:11.497804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T19:38:18.674355Z digest=sha256:61684ae8e18912ddb52f2080ef1c6a77b47479e6606a34bbd772e699a27d353a

Observation 0a7a6c42-8f3b-4ca0-8a9d-a7f391032f7d · inbound

Simulating the Evolution of Alignment and Values in Machine Intelligence cites this paper.

Simulating the Evolution of Alignment and Values in Machine Intelligence Large Language Models Often Know When They Are Being Evaluated

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:05:48.497134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T20:15:46.311347Z digest=sha256:1853435a4c69b6833ce7236402a158333e025f788ef255bfc2817122cd4bb903

Observation e931c58f-1d37-4a12-8957-f3dacf5c1eee · inbound

Honeypot Protocol cites this paper.

Honeypot Protocol Large Language Models Often Know When They Are Being Evaluated

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:35:34.167631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:35:16.230357Z digest=sha256:cf2fab79c461b3c2bc03095a7ff949f63124499712b8fd89f9bde00f56bfb3e9

Observation a5726432-a35f-4568-9ac5-ff01aa23a448 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Large Language Models Often Know When They Are Being Evaluated

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:16.451642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:301a88c11748065fe6f6508e4271be0b01b49016df463892d56d72aa50dd55b9

Observation 11820b76-29ca-487c-bbae-8368eaea64ef · inbound

Towards Understanding Specification Gaming in Reasoning Models cites this paper.

Towards Understanding Specification Gaming in Reasoning Models Large Language Models Often Know When They Are Being Evaluated

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:31:09.801473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T16:24:52.564505Z digest=sha256:62d68b234caabe937bc757da48c1798b6eed1c51abcec422931fa92cacdbaaf4

Observation 3ffaa199-acf2-41bc-b72f-3b1994ed3d86 · inbound

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use cites this paper.

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use Large Language Models Often Know When They Are Being Evaluated

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.651126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T15:38:41.821264Z digest=sha256:a2df11425ec57241e5b86bd284008f71810c861074b2b0112ced67f2f0a49054

Observation ec45d5b5-4ded-4809-9c16-e34ed3515f9b · inbound

Evaluation Awareness in Language Models Has Limited Effect on Behaviour cites this paper.

Evaluation Awareness in Language Models Has Limited Effect on Behaviour Large Language Models Often Know When They Are Being Evaluated

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:46:15.167301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T11:00:54.568772Z digest=sha256:60a14e3387ff82b99962a9225fb5245b674888000769775aadbd554ac3e4f2b5

Observation b4e8bf0d-15fc-494a-9c0d-aee4dc2e0600 · inbound

Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors cites this paper.

Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors Large Language Models Often Know When They Are Being Evaluated

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:10.915339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T09:44:44.131088Z digest=sha256:a9a56986f90395f42ec5ff225eeb9509f0a8316834c56662a17490dc40ab6cd6

Observation 2c504cd9-7b1a-42f8-a979-2cb3ec681887 · inbound

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels cites this paper.

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels Large Language Models Often Know When They Are Being Evaluated

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T21:39:24.666799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T12:07:02.778631Z digest=sha256:40a43cbf76dfcd865edd62a1527db28652914a8fa41ed84815a5bd7a818e9d74

Observation 29f87113-88d6-47b3-a880-7cfefb06b50d · inbound

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime cites this paper.

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime Large Language Models Often Know When They Are Being Evaluated

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:26.489069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:37:12.567351Z digest=sha256:f969434e5b53ff6111860758046c6d87a03a82158292676160d5fcb93fe1a57d

Observation d12d4985-7841-4d48-aacc-95684c94c188 · inbound

Naturalistic measure of social norms alignment cites this paper.

Naturalistic measure of social norms alignment Large Language Models Often Know When They Are Being Evaluated

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.547338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-25T04:36:48.254343Z digest=sha256:a388f20118e73a155b1b06ef152d07064cd26b4a37d24851e57cc87b73a43340

Observation 4f8a1ee2-2cfa-4ed0-b5cb-48c5b5f25f0c · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Large Language Models Often Know When They Are Being Evaluated

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:21.934177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:cd3f447e5462a96375ee112593935977dbc8b460ea8ebad1aaeb1c9c5da22175

Observation adfc2b2e-db10-4ab5-be71-542087ee551f · inbound

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents cites this paper.

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents Large Language Models Often Know When They Are Being Evaluated

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:36:52.855321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T04:58:10.803420Z digest=sha256:2e5fed6c08b2e65f8eab325e42ff210a7fab9db3495294e16e9e8e9f856e114c

Observation 6a8b9bfd-f0d7-44cf-8261-44acad1fd5d7 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Large Language Models Often Know When They Are Being Evaluated

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T01:41:29.315927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:110cc330720b5947101d1db60025c72c87958674a9333cac5a53457c09f272f4

Observation 5eb1a384-aa8d-4722-9772-0ed7bf494874 · inbound

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning cites this paper.

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning Large Language Models Often Know When They Are Being Evaluated

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.886101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T17:37:51.505359Z digest=sha256:2033151421c2783158c5932438a9767be1cb2d45ec853fe8607318e824cde613

Observation 62bfab90-7019-4ae2-ab7c-41c8a09278d0 · inbound

Sycophancy Towards Researchers Drives Performative Misalignment cites this paper.

Sycophancy Towards Researchers Drives Performative Misalignment Large Language Models Often Know When They Are Being Evaluated

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:27:26.195598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T18:52:29.375027Z digest=sha256:ad585ca7c8b90ee78e0d9c984cb1d91da7e5f58885713972418722ac3a368739

Observation 09e0708c-53d6-47e4-906c-e6fc347d0fe3 · inbound

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs cites this paper.

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs Large Language Models Often Know When They Are Being Evaluated

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:07:39.428576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:23:28.061924Z digest=sha256:9d6ed0bcdfd49594bace565173386ba7c2f50dc82312ccae278737aa507c81f0

Observation 37fd511b-e396-443e-ae66-40f8d5f092be · inbound

Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization cites this paper.

Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalization Large Language Models Often Know When They Are Being Evaluated

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:17:44.938579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T10:59:09.372559Z digest=sha256:31322b55a2e7c9ee6e568aaacb286206ef25f9c755a27091afe4cd4f573bf508

Observation 24a2c7fb-6793-4a10-95aa-d4e0f3d12700 · inbound

Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs cites this paper.

Normative Robustness as a Frontier for Non-Verifiable Reasoning in LLMs Large Language Models Often Know When They Are Being Evaluated

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.440185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T09:50:24.891865Z digest=sha256:42e40ed1ec6a5a08586b0792bf141dfe2bf8ebb242da8de8c5fcc195aa37eff2

Observation 512610eb-fc91-4a50-b0f6-ce0079167855 · inbound

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models cites this paper.

Evaluation Awareness Is Not One Capability: Evidence from Open Language Models Large Language Models Often Know When They Are Being Evaluated

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.869656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T08:23:20.122338Z digest=sha256:1067a9c00c90489cb9576e4529e1863cb70803709d270b39d63743a0885900b7

Observation 3242f6d5-421a-45a4-9aad-0d419c6ac36e · inbound

Defeat Devices in AI Systems cites this paper.

Defeat Devices in AI Systems Large Language Models Often Know When They Are Being Evaluated

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:44:27.919110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T08:34:58.879346Z digest=sha256:349e5389d23d7099629e843b62222bbaa5aacba21df59f47a62cfbc1839028dd

Observation 71ad9311-cc8c-4b1f-a7cc-ac1b262d95d4 · inbound

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action cites this paper.

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Large Language Models Often Know When They Are Being Evaluated

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.698122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T05:40:54.002702Z digest=sha256:cfb5267240e775cde52aff22387ec692cb87453e4655e97af2a97b91c2253c9b

Observation a59c605b-f5b8-4386-9a4f-7586d3f74aab · inbound

Predicting LLM Safety Before Release by Simulating Deployment cites this paper.

Predicting LLM Safety Before Release by Simulating Deployment Large Language Models Often Know When They Are Being Evaluated

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T18:16:25.813826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-09T18:07:42.556329Z digest=sha256:7ca2b4c5e780fa8497556073fbc0c9b4e513819ad0685a06dda26cfc6a72821e

Observation 3f181c66-37a7-4bc2-b6f4-0822b3871b4b · inbound

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation cites this paper.

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Large Language Models Often Know When They Are Being Evaluated

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:50.626336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:29:50.626336Z digest=sha256:4c09eba4cd777c8c66dff63cc49ac5183694c6017d1d98eaf16a144deef4acb1

Observation ba0632d1-31cc-4481-97af-30d0b63d7694 · inbound

Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models cites this paper.

Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models Large Language Models Often Know When They Are Being Evaluated

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T14:26:49.433333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:26:49.433333Z digest=sha256:d842447649dacc7c25c4ef040a0f43310a9b134281fa9ecf00f0ee514340fbb2

Observation 73131b97-d4f1-457a-844c-2339afd33228 · inbound

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory cites this paper.

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory Large Language Models Often Know When They Are Being Evaluated

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:43.718137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:43.718137Z digest=sha256:8e7b3034b639a170fa43a19c665eab9453b2160d0b44072218650f7469480d82

Observation b0cd2c7f-7416-43b1-b555-412c771768e6 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Large Language Models Often Know When They Are Being Evaluated

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:41.785644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:41.785644Z digest=sha256:6cd7ca92889382d92ec02e8e26c07a5473b845940be2f7c7e665a0096f183d22

Observation 592725c1-a39b-45e3-b101-7b47c9da458d · inbound

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models cites this paper.

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models Large Language Models Often Know When They Are Being Evaluated

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T01:13:36.803652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:13:36.803652Z digest=sha256:f3de949fecf622bbd93ccc0c4d7a3253ae921d6277a8b42d1b21a42211f5fb4c

Observation e6e48d91-61f2-4ac3-9bcb-31e893026e59 · inbound

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? cites this paper.

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? Large Language Models Often Know When They Are Being Evaluated

Reference 168

Resolution
unresolved
no resolver link, observed 2026-07-30T16:01:42.510478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T16:01:42.510478Z digest=sha256:dd86aada2697bed7620ae39aa5b11a0bcac41b77b7514d705938e6c6d4bb42a6

Observation 61b0ff3f-f4ab-4b02-87d9-3ba13e643245 · inbound

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures cites this paper.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Large Language Models Often Know When They Are Being Evaluated

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T00:25:12.383258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:25:12.383258Z digest=sha256:d08dc49a6747070876ce066a5ef5ef31b1f1fad3c473ee15e196626a50e662dc

Observation 1bd384b3-2bfc-4a56-898a-426a4397d627 · inbound

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation cites this paper.

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation Large Language Models Often Know When They Are Being Evaluated

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:00.896094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T17:02:00.896094Z digest=sha256:482cc0589f52cdfb114bb2f0f7dd518e6727df4c22611cda5ce77b5968268618

Observation 22b4dd25-ddf0-4615-be23-e88bd3efb9f5 · inbound

Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance cites this paper.

Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance Large Language Models Often Know When They Are Being Evaluated

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:19:26.030841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:19:26.030841Z digest=sha256:c87212b5837b4b5ac375867d4fa78469a32841263db5252947fba6fb15c6f9bc

Observation 3e85087d-387f-4c75-89db-e0fb1db3f69d · inbound

Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance cites this paper.

Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance Large Language Models Often Know When They Are Being Evaluated

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:20:32.813026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:20:32.813026Z digest=sha256:070a16913717a85112b40ac3607d203f5088adaabbea66a5d919fb00a0c458f8

Observation 141c6d5b-da9a-4c0f-96ff-bb6c0d23d9a0 · inbound

Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation cites this paper.

Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation Large Language Models Often Know When They Are Being Evaluated

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:00:07.878819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:00:07.878819Z digest=sha256:64c452639e514b65911a8c39ed20bc6f065bb7788ede023f12247bf1dc700aab

Observation 380638a4-fddc-4bc8-858d-53a9ee46c614 · inbound

A Probe Direction Is a Property of Its Prompt cites this paper.

A Probe Direction Is a Property of Its Prompt Large Language Models Often Know When They Are Being Evaluated

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T13:38:50.766477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:38:50.766477Z digest=sha256:d9a3dfdb9e53e63fda391a0d504b82c2414c025dec6b1031594e4b884e11ba02