Pith. sign in

Paper Citation Record · LEDGER

Fundamental Limitations of Alignment in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2304.11082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.11082 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:54:36.557710Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

44
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc298fbf-da94-4833-8e0d-3f5871613c44 · inbound

Jailbroken: How Does LLM Safety Training Fail? cites this paper.

Jailbroken: How Does LLM Safety Training Fail? Fundamental Limitations of Alignment in Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:17:42.828873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T18:17:42.752997Z digest=sha256:ba5c294da4f3970f4caf461e1275dea19a3b860f3a25a605ecb3804d041a4c47

Observation 3efae1f4-54d7-4687-aafa-7b082e8c0467 · inbound

Universal and Transferable Adversarial Attacks on Aligned Language Models cites this paper.

Universal and Transferable Adversarial Attacks on Aligned Language Models Fundamental Limitations of Alignment in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.470910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:a940c7a02e0effc7e92a6969b1656f921c7a5fb1a041c8ac33750731e716b69e

Observation 6decc58e-6b91-4b12-af93-4bdc9bc1e436 · inbound

Hyperbolic Deep Learning for Foundation Models: A Survey cites this paper.

Hyperbolic Deep Learning for Foundation Models: A Survey Fundamental Limitations of Alignment in Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:36.557710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:36.557710Z digest=sha256:e7fd0b4ef58d5cba4ff9d6a343a83c8262c637802890289d964ac52eda0b068a

Observation 560af3c0-046c-40eb-90de-d44ddc6a6834 · inbound

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring cites this paper.

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring Fundamental Limitations of Alignment in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:51:03.796947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:51:03.796947Z digest=sha256:9fe39eff370b28f488657bbc9b09ace9eaae3546fccd85c55cddb60f40e87eba

Observation 5e0dd17a-3fe9-4cf7-97a3-67fc29cff162 · inbound

Robust AI Security and Alignment: A Sisyphean Endeavor? cites this paper.

Robust AI Security and Alignment: A Sisyphean Endeavor? Fundamental Limitations of Alignment in Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:58:38.228561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T22:58:29.306435Z digest=sha256:e7c95be692b1f744e6abb8e2de3271c9919599e6075fd8cf26dfea5ae6f33cc1

Observation 6a98c7df-c782-4da9-8ac3-581311a95014 · inbound

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations cites this paper.

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations Fundamental Limitations of Alignment in Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:18.723757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T06:04:24.077276Z digest=sha256:d73a8ff4a646f8f6b3be67067bc6d5a912889035a7fee1932e09a47615fdcc84

Observation 5d7d489b-4ec9-4eb0-8d06-8dab12957bd0 · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Fundamental Limitations of Alignment in Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:43.581200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:08b7e8a6e2b939fbf46081259aa743d10827db5dd986ea990180f6b3ab4db543

Observation a9fe88b3-3c82-4db0-9281-54d53f7f3910 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Fundamental Limitations of Alignment in Large Language Models

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:55.556966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:df9653f06f0822fbc90a63b5f3fc4228b287c349432dee91ed1adb81c683ee1b

Observation 9b99e65a-6ad6-4fdd-bcf4-65cbeffdebe5 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Fundamental Limitations of Alignment in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.130331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:09b1523ac3dc0124015499f5b9810276522ff51bcab36e2e2550775dbdadd7d1

Observation 3de5c74b-9fdd-42c4-8631-8a55023605de · inbound

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs cites this paper.

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs Fundamental Limitations of Alignment in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:49:35.636729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T04:45:35.079192Z digest=sha256:6a7b9bdb7d16d84cf3e8e2a0023ffa8635e05736edc60dd256439f119129a8c3

Observation e70bd3cf-98ea-48d2-84aa-478bb6a88973 · inbound

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible cites this paper.

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Fundamental Limitations of Alignment in Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.928075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:28:02.493124Z digest=sha256:71320043bd66aabb0f80452831b3209803fdb6866ffd5bdd6962d036678dd868

Observation 669a96b2-01be-4161-880d-11824b753b97 · inbound

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible cites this paper.

The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible Fundamental Limitations of Alignment in Large Language Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-02T13:17:57.695515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:17:57.695515Z digest=sha256:e4f322c852110af7ddd7bbb1c87fca0f971fdb5efe97d5b944dfb80e56169ccb

Observation 7fab3428-33eb-4396-bbc3-6788746fba32 · inbound

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms cites this paper.

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Fundamental Limitations of Alignment in Large Language Models

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:32:53.150041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T00:26:54.019256Z digest=sha256:bc919380e9e744cfc18feb9db49338b3b60005020cdec8deab752e16b47f4ef6

Observation 8fc14387-4f84-447c-8574-00d7ebe85e8d · inbound

Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs cites this paper.

Confused ChatGPT: Cross-App Context Poisoning via First-Party APIs Fundamental Limitations of Alignment in Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:35.659788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T18:55:45.465260Z digest=sha256:e87a760846e1a82538a403264e7c65f332cdf22a85e7894dcf8d1688a9b20958

Observation def98b02-acfb-4a38-bbab-17645d140d20 · inbound

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy cites this paper.

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy Fundamental Limitations of Alignment in Large Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.953535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T18:36:44.265273Z digest=sha256:165764587fb783e398061a9d96507920817efe60ed7d2fb16bf707c8665f30c4

Observation 92bd060d-0b28-4cbd-9348-c6f9896922f7 · inbound

On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study cites this paper.

On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study Fundamental Limitations of Alignment in Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:18:03.179730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T09:40:48.736006Z digest=sha256:63ed50b2491dbd04291c07951641f2e966c8afeffd2133a37bb832e15b215b37

Observation d459a3d0-42eb-458d-81b9-10a3b9baff94 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Fundamental Limitations of Alignment in Large Language Models

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:31.376035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:31.376035Z digest=sha256:1637b8a79507cd443a1e376459ef3f31c7fdf7e572220a230d9c0a19948fa2d4