Pith. sign in

Paper Citation Record · LEDGER

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

As of 16 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2507.02977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02977 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:22:59.700209Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:48:43.262187Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:45:45.914511Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14b7a6ad-30ac-4e50-a1a3-671cd2f4a38c · outbound

This paper cites Openai o1 system card (sep 2024), 2024.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Openai o1 system card (sep 2024), 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.888753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:58.393895Z digest=sha256:d49650c377d5ffa13dc284adca55d0b3be21f198e30cec4d01c24c26f8f020b7

Observation c839a1b3-e8fc-4d37-bd23-900c89b86489 · outbound

This paper cites Troy, Stuart J.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Troy, Stuart J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:01.486686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:58.455386Z digest=sha256:0b3dbc3f40c38a127a11037d7d55ded4e18acaef4cf76ad82aeed0f940a672d1

Observation b162460f-f4ce-4b8e-bf00-7117e17111b8 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Frontier Models are Capable of In-context Scheming

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.636617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.636617Z digest=sha256:4cab7efd402addc2bbec199fb71ab8e6aa73e12e636eb8d87a596cc032a9b8b6

Observation da5ee6ad-0100-470f-893a-0314c97069f4 · outbound

This paper cites Demonstrating specification gaming in reasoning models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Demonstrating specification gaming in reasoning models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.773703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.773703Z digest=sha256:07e70e0ee992c6705d6b3ee8a5f171e463a25af7b721bed9952ee73982c8eba5

Observation b9a3c403-56a8-4381-a94f-b4874e0c9b15 · outbound

This paper cites Alignment faking in large language models.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Alignment faking in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:58.816501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:58.816501Z digest=sha256:9c00838172cce8240a27c5d859e303384e4c1fcd386801af4e2293df770f00e4

Observation f84e754d-9041-4a9f-ab13-352b95172087 · outbound

This paper cites I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance I replicated the anthropic alignment faking experiment on other models, and they didn’t fake alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.987883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:58.923788Z digest=sha256:4fbb499b0c198aaa98b375907ac0585b12a1cfed3ffcdf6331f30038263231f8

Observation caddabb4-294c-48d2-b1e5-78e4f70d6b39 · outbound

This paper cites Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.013316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.013316Z digest=sha256:c2764dd4bf620ba9b727efab70bcf404b76d54d008ff1a65b2c86dd8f2461a79

Observation 948e0b54-c2a6-436d-bc62-e1fc37608a58 · outbound

This paper cites Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.018086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.018086Z digest=sha256:387f2ef86dbcc33f1d0cbb201ea08778e2c14fa6d351f89570d9c84ff7e4c205

Observation 0b532cbf-e85a-455d-b071-122038bcb9e8 · outbound

This paper cites LLM Agents can Autonomously Exploit One-day Vulnerabilities.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents can Autonomously Exploit One-day Vulnerabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.085926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.085926Z digest=sha256:6ea4ec52296e095bdf7effa3346a4acff3d5dfd20d765def68ad98314cf6aa42

Observation ac2932f3-7626-4f90-8d04-0933120e2807 · outbound

This paper cites LLM Agents Should Employ Security Principles.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance LLM Agents Should Employ Security Principles

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.187197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.187197Z digest=sha256:54aed7aceb42f90ac504c98d28207f08f9a0bacc7fba98923335321919b32068

Observation 864ad225-6ec8-4119-bf39-664983f92734 · outbound

This paper cites Gemini 2.5: Our most intelligent models are getting even better, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Gemini 2.5: Our most intelligent models are getting even better, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.764807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:59.246200Z digest=sha256:77431a743b4f5eadca57178520dbdf3e00ffc8d19863ae325517dda2b023a6c7

Observation f8715da1-640b-4c2a-b3d8-ff1984219aba · outbound

This paper cites o3 and o4-mini system card, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance o3 and o4-mini system card, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.588657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:59.363069Z digest=sha256:084fe82dcbade7eb182c1aff84ea61bc32d56c40d6935d9181db4b23d22bbf74

Observation a2429620-6956-4cc3-a50e-c31eb707128d · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance System card: Claude opus 4 & claude sonnet 4, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.394776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:59.442120Z digest=sha256:d225f923bb3dcb549108481ba644a156a2b1085606126d90b1e85832d66e6e0f

Observation 36056a68-c910-46cf-8ba1-87983f700e77 · outbound

This paper cites Deepseek-r1-0528 release, 2025.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Deepseek-r1-0528 release, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.214313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:59.518923Z digest=sha256:d6067d8a3aa4439c0ba04d0d8b918c981763aa0ba9bc1a037153bfadc322b4e7

Observation a75e8976-5b5f-4ce4-9f19-0589630e90d9 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:59.603290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:59.603290Z digest=sha256:f0b87dba09b60934063ace4fcc5a7b343670494efc80235d1ce24e9130d2dab3

Observation e0e97d17-b714-42de-897c-4af5f0637f63 · outbound

This paper cites reference.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance reference

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:00.014538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:59.700209Z digest=sha256:6a339b561f174a87236acec7281130a6fe46eac6856b4cbdb9ebe8fba989a29b

Observation 5497942e-b2b9-4158-b589-003931fee4ad · outbound

This paper cites an unresolved cited work.

LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:01.201836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:22:58.506406Z digest=sha256:8ed1e7a7081ad0440fc9999918c50c4cb6fd92cd837730233d66466bff07b86e

Pith citing papers

Observation 5ab2a844-f601-40cf-b5ac-f63a4e8c73a8 · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:36:24.427998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T05:33:40.813795Z digest=sha256:12b8a85502b1d676edc9389bc34773f3c29c6ab715251333eabdb05a2479b625

Observation c237a8e7-8289-4fec-8ef3-ab1df0f3d69d · inbound

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems cites this paper.

SciIntegrity-Bench: A Benchmark for Evaluating Academic Integrity in AI Scientist Systems LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.916404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:48:43.262187Z digest=sha256:b835bd073155e5987da150e6aa597f21bc9817ab7d0cf9308266d43dc9093a88