Pith. sign in

Paper Citation Record · LEDGER

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 54 inbound Pith citation observations for arXiv:2404.09932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.09932 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:53:20.428520Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e47d2fa6-25b8-4557-8bcc-89e4f3513b11 · inbound

Scaling and renormalization in high-dimensional regression cites this paper.

Scaling and renormalization in high-dimensional regression Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T01:55:55.089386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T01:54:48.781227Z digest=sha256:6dec72cb2f01364301c3df7ed7174371e22422d87505133fc5f304b305bfd52a

Observation 07a1da5b-b4f0-44e8-b987-1a99bf9035de · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T16:25:14.801617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:411d4ddd414810d9626690f2b701673d9fd274cad2dcc57fe9434b725efc2f33

Observation 9c9485fd-3d9b-48a2-a535-f7d52dd85d64 · inbound

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense cites this paper.

A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:53:20.428520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:53:20.428520Z digest=sha256:14830ca7475bebc4bfd948e4bfa299e873f196c32881293cc97ce50b364da091

Observation b11812df-a75e-499f-a9e1-35055655b0f2 · inbound

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs cites this paper.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.875258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.875258Z digest=sha256:3e1246841a09bef1a338e896b8e1d57b95e1c80d67d7756fcb84ab24e6fabbeb

Observation edaf3004-9b57-4970-a514-98992a9076e3 · inbound

Mechanistic understanding and validation of large AI models with SemanticLens cites this paper.

Mechanistic understanding and validation of large AI models with SemanticLens Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:11.730920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:11.730920Z digest=sha256:f516f37b5bcd18a601d0bd1837458e889e11bb23b4a9c54341572774ca4f8c6f

Observation 6bb033e1-ce8e-48c9-a83d-809d6234bd9a · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.714194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.714194Z digest=sha256:b0a1e0d6e226ba658a546ea1eb0edd9d1868ccab447de2c42b9f6add3a386bc8

Observation 4a0ec520-0d0d-45c1-97ad-35ebf5e33ec5 · inbound

Clone-Robust AI Alignment cites this paper.

Clone-Robust AI Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:11.764071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:11.764071Z digest=sha256:ca81c0d38b9301ce6bf01e972346be7a047ab2dc53569275178bdb1b8015e76d

Observation e88b23e9-baaa-4345-a6a3-3101215e6698 · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.488108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.488108Z digest=sha256:63bbfd91a5a196412100447c452656203dddffbffbd89bd82e0d5909b284fa20

Observation 9a148bf4-fa45-45ed-9888-668ed4c51d37 · inbound

Episodic memory in AI agents poses risks that should be studied and mitigated cites this paper.

Episodic memory in AI agents poses risks that should be studied and mitigated Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T17:58:17.922278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:58:17.922278Z digest=sha256:055294e322a0eed6a59817975ff0c75c7c37866766de20d09bb7450495303887

Observation 926cd96f-cc6f-4f4b-ae9a-489852782b6f · inbound

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models cites this paper.

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T14:51:58.699740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:51:58.699740Z digest=sha256:4d3e6ab6e1b8aa87070742c70b80520fe0e5e44d5e43cee4aa2c7fdc26d48845

Observation 2e708e35-c5da-4e6d-831e-ba710e2645fa · inbound

The AI Agent Index cites this paper.

The AI Agent Index Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:48:34.657585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:48:34.657585Z digest=sha256:efabcbe94fe3c85e1d3cfbc4aa997bb63dbc40855f42eb53565238e0ea6f96d0

Observation 0d821892-22fe-43fa-9cf9-a296911e69e9 · inbound

MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf cites this paper.

MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:09:16.102465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:09:16.102465Z digest=sha256:50dee1d6d9b17b59711e807472f82d7eda3caf037bc78f01e68280058b8d6020

Observation 67090237-d39b-4ec6-85bb-953388fac2cc · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.511215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.511215Z digest=sha256:b97c4ecba971dcd3f3fe332b12288c322f62092f80329b934ec8e5ea19cb211a

Observation 276b0fa3-77a0-4277-a525-fee2aa77b9ef · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:20.088578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:20.088578Z digest=sha256:99275a2e731aa2f03e34070077cd3140976ba8f022c920f9c15a7bdf92886e6d

Observation 0c2bc9d3-9b3a-4bd8-a2df-7f37c0a94918 · inbound

Mitigating Deceptive Alignment via Self-Monitoring cites this paper.

Mitigating Deceptive Alignment via Self-Monitoring Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.428012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.428012Z digest=sha256:d4f3691d0381df78409bcf195e8a7f3503a7c8d7f96bf3c4e9ebd70a9be00955

Observation eaaae2ea-aa11-4887-9fc5-b7df25e0b9b0 · inbound

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment cites this paper.

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:10:51.554189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T01:06:19.756032Z digest=sha256:09254dc8fd61cae0a9834d944f28df0d6fdca65193f6e9a6cd305b7c59cc57fb

Observation 6604e4c7-4876-4fbf-a6c5-4753b52ba655 · inbound

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models cites this paper.

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:30.354244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:30.354244Z digest=sha256:8e8ebe76841b713dbd4b29395b788d3dbed751fc1a2c484cdd3ea8c318b9aff0

Observation 0d1cbc67-e62a-4ab9-8b73-d7ef4d75dfb6 · inbound

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It cites this paper.

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:13.874731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:13.874731Z digest=sha256:9901a98b403ac44fd3ed5b21a0d29f2e48a47b066de4eeb79127e7813edc3467

Observation 75c56f71-475a-42c5-8fca-4ab6831a08d1 · inbound

Risks of AI-driven product development and strategies for their mitigation cites this paper.

Risks of AI-driven product development and strategies for their mitigation Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:58.104897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:58.104897Z digest=sha256:4566c5b6bebda372a15b7c70c718d622b61aec60e462882906e42b05e9cd7eac

Observation 6ffa341c-45b3-4e5d-9adc-c2fa387f1b01 · inbound

Linear Spatial World Models Emerge in Large Language Models cites this paper.

Linear Spatial World Models Emerge in Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:28.039432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:28.039432Z digest=sha256:8f7da9baacea42319906bb347d72f489cb1c7af859d8c7cbc090da4def562138

Observation e14fbc27-e4c3-4582-9af6-2e467cb1714f · inbound

AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents cites this paper.

AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:57.913128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:57.913128Z digest=sha256:3483d5b14f75ef14f80e8eeda0194bb93319c49fa90d58ce1405927ed6577a60

Observation e62633f9-e83d-4a72-b74f-41265a3c4fd3 · inbound

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems cites this paper.

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:36.451035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:32:36.451035Z digest=sha256:86ffa872444dbf1d303ded7bef9eef6a4720ef946e84030cd405bfae48f2ebd3

Observation 28c35c6b-b441-4b2e-a413-138b5172816b · inbound

Probing the Robustness of Large Language Models Safety to Latent Perturbations cites this paper.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.094114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.094114Z digest=sha256:5aa64f71a5dbfba3f9eb40883c9498211b69aa258ca6cc12cde19405fe5d4365

Observation 555ad9b2-e220-459a-b447-9abe1b5f31e0 · inbound

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models cites this paper.

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:22.466382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:22.466382Z digest=sha256:50140bf364e54771ee34fa236eab2c62cf4057d8359f5dccda098e348e0233ce

Observation d718aa0a-df81-4905-89ce-89b731533a5d · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.377039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.377039Z digest=sha256:d2ba11e6fabafda71c3d7753ab0113c00eb7b94a73f4530a5a9450057eb0da76

Observation 3e01fb38-ff2f-4f15-9450-db40c886c0bf · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 292

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.423464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:d5cd1b4084f12e30b86283c6d997d2ca54f05fcad7d44e3717b79916fc81d00e

Observation 2e208163-7343-40e9-b35a-595ddf22ab0f · inbound

Against racing to AGI: Cooperation, deterrence, and catastrophic risks cites this paper.

Against racing to AGI: Cooperation, deterrence, and catastrophic risks Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:22:27.592346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:22:27.592346Z digest=sha256:2a25f426249c5d53d5650f76fbbfc4cbb8aa155bb31b3063920f92da795aed82

Observation 372bda71-ed0d-4395-82e4-c3299076289b · inbound

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions cites this paper.

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:24.795223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:24.795223Z digest=sha256:2d10e572c57eec4248e309b9b2170f66209b871213983e67a97d7c5c71a61da2

Observation e0f6b72d-cb2b-44a1-a86b-39d5e21a3c38 · inbound

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information cites this paper.

Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:07:02.902944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:07:02.902944Z digest=sha256:7c2ccf986deeabe689606fc0b263f4f76e79e74eb776baadea3532d732fa4805

Observation ecc3fec7-e818-46c6-b563-51e46ebf87d9 · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:31.598124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:31.598124Z digest=sha256:ba21d889c453d09a6560754209272cd739da610420d1d2a434e9eaa1ae05dc00

Observation e8c68cc1-77b3-4e49-af8b-3d246cf85831 · inbound

Scheming Ability in LLM-to-LLM Strategic Interactions cites this paper.

Scheming Ability in LLM-to-LLM Strategic Interactions Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:51:03.750005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T07:50:30.597108Z digest=sha256:fbe10f1868a5dabe7266e8d77456e4265251087a807931bf48e002fa086acaef

Observation c691439b-6c02-4607-97ce-bce0e5225742 · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.304935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.304935Z digest=sha256:269be5686782adc394a0d696912a8ec42d674b03e17bffd84dbfdccd142f5302

Observation 6777e281-4e8f-406e-9a6a-71c8f5b6d78c · inbound

Phantom Transfer: Data Poisoning can Survive Data-Level Defences cites this paper.

Phantom Transfer: Data Poisoning can Survive Data-Level Defences Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T04:59:05.071079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:59:05.071079Z digest=sha256:b2adc9e912ebd0eee086dc0c3b35841a5a044d5456872199412080628f247501

Observation 41cc4a35-8e10-49c1-a665-dadfb8cdaed7 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T11:21:28.998865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:7a8ec77eabb9a88d644dc5da93f4ddbcdca1562348abefd8e9582c4336de3922

Observation 4bc4ceda-8734-4625-addb-79961c6c573b · inbound

NEST: Nascent Encoded Steganographic Thoughts cites this paper.

NEST: Nascent Encoded Steganographic Thoughts Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T23:21:31.165458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:21:31.165458Z digest=sha256:e00b0cc0f01e51a40ce16402d56f3da376eeab70787a7a4afa437488f3aacdbf

Observation 21fcd595-f29d-4fd0-8931-13f04dc59457 · inbound

BarrierSteer: LLM Safety via Learning Barrier Steering cites this paper.

BarrierSteer: LLM Safety via Learning Barrier Steering Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:05:26.712893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T07:02:03.058731Z digest=sha256:360df0fb6c9b6c052aa89a405385ff646e4bfb50dc40e9df5f665efb7b72f144

Observation f1f45249-5e71-4b83-be97-9d66264b82b8 · inbound

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior cites this paper.

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:16:06.543864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T17:47:09.591001Z digest=sha256:c88bfa5a4c24bf6aba75b8726b4d3f8ee456f0bcba0206db1d0811c4e361c897

Observation ffd801da-ca43-4ed6-b6f6-99f976febded · inbound

Belief or Circuitry? Causal Evidence for In-Context Graph Learning cites this paper.

Belief or Circuitry? Causal Evidence for In-Context Graph Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.962333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:52:07.392800Z digest=sha256:b9e3e9f5a0a861ef8d6c7ba01abf2bd17cfc8d17328ee9fad90e7278ecac7d46

Observation 2bcdda77-c8c6-4b91-8ba4-782e7a398082 · inbound

Interpretability Can Be Actionable cites this paper.

Interpretability Can Be Actionable Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:17:22.798229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T06:12:51.656452Z digest=sha256:284586f383d0af8e8fdcd5d22206205cb339b946225374f683e904f9cea3230f

Observation cf4c3b87-5ed5-45d5-b964-48f9624777db · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 144

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:19.282765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:fa4f0ece5640d9c8bab07e458a92230a986ab21b6c4d960020fd0dce00704f87

Observation e4284353-edd1-4007-8b40-4521a20c3b61 · inbound

Prediction-Powered Inference Across Many Tasks for AI Evaluation & Social Science Research cites this paper.

Prediction-Powered Inference Across Many Tasks for AI Evaluation & Social Science Research Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T06:03:08.437075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T06:00:59.614364Z digest=sha256:0c5181da82e3a88d3a8fc09232e1e358537391e32965cd21b26a9e94bd5a748c

Observation b7dfc7f5-81bb-491b-ae4d-bf9f74b8ffa8 · inbound

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms cites this paper.

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:32:53.047592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T00:26:54.019256Z digest=sha256:e8b631ebd6d0b630e20997df80c6758616308a015dbd4f9ac6e34c64f8a74909

Observation ca044699-010f-4031-944a-97e2aeaea091 · inbound

The Surface You Test Is Not the Surface That Breaks cites this paper.

The Surface You Test Is Not the Surface That Breaks Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.190683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T06:37:19.674012Z digest=sha256:24be78d8c9a25ab29346fc0110878ad690ef327d25b3ffada7e986e4ee4f3583

Observation ee43106b-6bdd-435c-b545-49b57857f71f · inbound

VET: A Framework for Analyzing AI Discourse cites this paper.

VET: A Framework for Analyzing AI Discourse Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.479301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T14:41:17.714421Z digest=sha256:f115a7f243567d0e5ce74a81bc136943a219ba06953b12b623c7b8ce540fc37c

Observation fe12e385-1113-407c-b812-b8418590317b · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:32:18.165223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:dbb56722f8969b52589f5460cb32b6671281d9d0af936b2a058eda7658a24798

Observation 2e78d5e7-5ce1-4849-b17c-6a011716ecbf · inbound

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs cites this paper.

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:16:34.761998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T09:21:57.373862Z digest=sha256:54ed8da0accdf5ab5970400080a184b88675dbe373b218507d696797b77fbe6d

Observation 46a39709-f2b5-41aa-892b-53ffb6583773 · inbound

A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders cites this paper.

A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:47:09.934660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T22:22:50.474397Z digest=sha256:55c9812a41bc9fb4ed831ac8341d1ba92670737061e38c128e9d63e55fe3f7bc

Observation dcbdf9e3-9228-4c45-bb78-1a1947a62269 · inbound

Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard cites this paper.

Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.017226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:04:07.735560Z digest=sha256:5c4f6334d05f01a54b5e76d6d7db1d4e21c9943f4694e7182c5ab4049b9a540e

Observation 5eb8546a-7b7b-4288-9804-0fd91738197e · inbound

Distilling Safe LLM Systems via Soft Prompts for On Device Settings cites this paper.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.222546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T17:15:51.375580Z digest=sha256:9825f325ddd5ef19492aa4bb63dc90ddff572a8734aaa3cdeb216436ef63892f

Observation e2cc9341-e13b-4344-aebc-f97c301d251d · inbound

Investigating The Security of Modern AI and Cloud Infrastructure cites this paper.

Investigating The Security of Modern AI and Cloud Infrastructure Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:29:42.338587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T11:31:39.910784Z digest=sha256:8a1ed461a166d88239279abb247f79ef7a1a725446a4b36dbaea0b3becc308fc

Observation 2811f24f-910a-4c3e-b60d-9ea0d54caf92 · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T22:40:37.839133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:40:37.839133Z digest=sha256:297e67aaf3c19bab51758cdb12d85b6ea7db16e5a88ba1e2a5e8f0cf7723975f

Observation 580bb6d7-6b73-421a-b14f-02db6794d56d · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T07:01:49.222325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:01:49.222325Z digest=sha256:f006dc66edf419cb022e296a8b25d03616902f6ad28a197ea9e9388331b557f3

Observation 078c6946-2dbb-4e66-bc77-84bc0ee45a14 · inbound

(Towards) Scalable Reliable Automated Evaluation with Large Language Models cites this paper.

(Towards) Scalable Reliable Automated Evaluation with Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T12:20:07.096313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T12:20:07.096313Z digest=sha256:93baabd36551544cf77cf0729ab63772bb8c3533cc1537f326cf4835a27ff68c

Observation b30a7b2d-a07f-4801-bbb5-93ca6fd4e08b · inbound

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools cites this paper.

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:11.241994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:11.241994Z digest=sha256:459276cb263c7dda1a189d173d742fd35795314c10cbb4d6b4755372b89e9a9b