Pith. sign in

Paper Citation Record · LEDGER

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2507.02977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02977 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:22:59.700209Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:48:43.262187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:45:45.914511Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14b7a6ad-30ac-4e50-a1a3-671cd2f4a38c · outbound

This paper cites Openai o1 system card (sep 2024), 2024.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Openai o1 system card (sep 2024), 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.888753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:58.393895Z digest=sha256:6e8319403fe34200fdc3c22eebc5af5c9ee6b113ff99f9e9706a246cc3443138

Observation c839a1b3-e8fc-4d37-bd23-900c89b86489 · outbound

This paper cites Troy, Stuart J.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Troy, Stuart J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.486686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:58.455386Z digest=sha256:1183c90bffabed095e823ed6a6a29a939a83a8027021c4af17ce50500fbf8f91

Observation b162460f-f4ce-4b8e-bf00-7117e17111b8 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Frontier Models are Capable of In-context Scheming

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.636617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.636617Z digest=sha256:b1b7b6501aacd1c3b26d0f7ebb1b1eb633e9597bb01fde821b94f5d0d838c417

Observation da5ee6ad-0100-470f-893a-0314c97069f4 · outbound

This paper cites Demonstrating specification gaming in reasoning models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Demonstrating specification gaming in reasoning models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.773703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.773703Z digest=sha256:1b8a289415278653c929d3a3e5a189d80398859dd77fabfb97c952c8c6d6fc61

Observation b9a3c403-56a8-4381-a94f-b4874e0c9b15 · outbound

This paper cites Alignment faking in large language models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Alignment faking in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.816501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.816501Z digest=sha256:b34af08ec66657ccfb0de658e53a0f4de2aad7629d5b989a7944dccfa2561611

Observation f84e754d-9041-4a9f-ab13-352b95172087 · outbound

This paper cites I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.987883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:58.923788Z digest=sha256:96dc4b4e53396bbebcc1c8a8a3bb37b67d1637ccf937aef8bb30ebf0d6902c16

Observation caddabb4-294c-48d2-b1e5-78e4f70d6b39 · outbound

This paper cites Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.013316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.013316Z digest=sha256:952a5e03c27dd868acebec29137f906ba73123b8c23bfc53ec3e89c88efc097c

Observation 948e0b54-c2a6-436d-bc62-e1fc37608a58 · outbound

This paper cites Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.018086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.018086Z digest=sha256:1987ad26e8eb3dda5878d934999f2e4263bef61e25b9f6e327fb41e385558f58

Observation 0b532cbf-e85a-455d-b071-122038bcb9e8 · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.085926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.085926Z digest=sha256:d8d001edbf4ccf4c7b39ce8a88dc19dbc5b258f9b8fe08d095c840f84c20b9a4

Observation ac2932f3-7626-4f90-8d04-0933120e2807 · outbound

This paper cites LLM Agents Should Employ Security Principles.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents Should Employ Security Principles

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.187197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.187197Z digest=sha256:395ce19947f83c3542560d899a6e54a778c0beded9f90845e184e47e0c1660e8

Observation 864ad225-6ec8-4119-bf39-664983f92734 · outbound

This paper cites Gemini 2.5: Our most intelligent models are getting even better, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Gemini 2.5: Our most intelligent models are getting even better, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.764807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:59.246200Z digest=sha256:a4658cc7e4889b356c01b9e17a3c05293ee887d9bf6e7a69d46b5bde4380e1ce

Observation f8715da1-640b-4c2a-b3d8-ff1984219aba · outbound

This paper cites o3 and o4-mini system card, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance o3 and o4-mini system card, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.588657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:59.363069Z digest=sha256:96022dd2b622cb8506b237fb12a547b5b99a2e08f5e3b524c8f33717ac0b1ca2

Observation a2429620-6956-4cc3-a50e-c31eb707128d · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance System card: Claude opus 4 & claude sonnet 4, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.394776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:59.442120Z digest=sha256:231522fec5cb03e0f1c88570390c6f2a0aa15cce2cce387fbfea68be3acb0e6a

Observation 36056a68-c910-46cf-8ba1-87983f700e77 · outbound

This paper cites Deepseek-r1-0528 release, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Deepseek-r1-0528 release, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.214313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:59.518923Z digest=sha256:0d938420a41dfa54181e587aeceae7124a4972bd4f13c70e7286ff4f241a1afb

Observation a75e8976-5b5f-4ce4-9f19-0589630e90d9 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.603290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.603290Z digest=sha256:4eb2a4610fba7a7860dabff74147db6ccb4a29e19a6621c5de8d427411bf6028

Observation e0e97d17-b714-42de-897c-4af5f0637f63 · outbound

This paper cites reference.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance reference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.014538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:59.700209Z digest=sha256:5ced2a487d186ed842beef63719a67a2d0f1ed86214b268813bee0017ba39f23

Observation 5497942e-b2b9-4158-b589-003931fee4ad · outbound

This paper cites an unresolved cited work.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:01.201836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:22:58.506406Z digest=sha256:16c2694e3cc4626d9df9fa32996771fec4d5f7666e8a6fe13cb6c2019015e143

Pith citing papers

Observation 5ab2a844-f601-40cf-b5ac-f63a4e8c73a8 · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.427998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:33:40.813795Z digest=sha256:03701c01f85d5bbb36887568b8097204a8732698662ba9dd65e6ae1ca6f33f2f

Observation c237a8e7-8289-4fec-8ef3-ab1df0f3d69d · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.916404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:48:43.262187Z digest=sha256:4d23d993620251536179e12841aedaf463162437adf4460abb54f980166664e1