Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:07.337382Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2506.13774.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:07.337382Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 03ffc49a-1a7b-429c-8020-54a96502feb0 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Artificial Intelligence, Values, and Alignment
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93026cbf-d874-413a-9bd0-b2651e5c8a3b · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Artificial Morality: Top -down, Bottom-up, and Hybrid Approaches
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e4ec19f7-68a5-4123-90ff-cb5f9267b597 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Translating Principles into Practices of Digital Ethics: Five Risks of Being Unethical
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cab35d18-a447-4352-ac03-fd37e5153ac7 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5354f90-fad6-4bb2-96cf-a66ca53ebf2d · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Personalized Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e8007a-9ada-423f-8dbc-f1bc1335f848 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Towards an End -to-End Personal Fine -Tuning Framework for AI Value Alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a1f78d6b-2a43-4697-b583-f52683d34c35 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Safer Agentic AI
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab4fc287-4bcb-4053-b2df-466b9746a33d · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Introducing the Model Context Protocol
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0b0b6157-c207-4c1b-b451-66d573f3884b · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Deep Reinforcement Learning from Human Preferences
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c1a6a662-4abe-47f4-b593-49c0d6ede551 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Confabulation: The Surprising Value of Large Language Model Hallucinations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f881bc-d2a2-4cc7-a983-2f2387095612 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Choice Vectors: Streamlining Personal AI Alignment Through Binary Selection
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8843945-5b73-4fb6-8fdd-6eacf9e521a7 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 232768fa-1ba1-48b4-b37f-24a77dad24c0 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Universality of Representation in Biologic al and Artificial Neural Networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5015ad3-26fb-48f4-8a6f-ff1489e4c7a9 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The neural bases of cognitive conflict and control in mo ral judgment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f808b93-0ce8-402b-a21e-fc5f0480cb73 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The neural basis of human social values: Evidence from functional MRI
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2d831fb6-b9b9-40aa-af80-9648e02cea1c · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values A Cognitive Theory of Consciousness: The Workspace of the Mind; Cambridge University Press: Cambridge, UK, 1988
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6886e3ee-5210-46b2-b28e-efef9e830027 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unified Theories of Cognition; Harvard University Press: Cambridge, MA, USA, 1990
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bec20e91-f154-4787-a276-9d29e7bc19d6 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Revealing economic facts: LLMs know more than they say
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 565c71ff-9edc-481f-a63f-4b03b833470d · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values ShieldGemma 2: Robust and Tractable Image Content Moderation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c700b9-bd04-4265-a0d1-417a7d2802c3 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Superego -Agent LGDemo (Branch: Fastapi_Mcp)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c55222a1-896d-459a-9f3e-dfd48bfa9033 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 218dadac-60a5-4f85-909e-cf5e6a25d19e · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values AgentHarm: A benchmark for measuring harmfulness of LLM agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d21dcae-dcab-4951-b7a8-6e0cf673c985 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Do the rewards justify the means? Measuring trade -offs between rewards and ethical behavior in the Machiavelli benchmark
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7d230248-3583-4771-9434-10d2f6bcc2e5 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a1bd80-b170-4952-ac58-bb35175f738c · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Vijil Test Library: Evaluating LLM Trustworthiness Across Eight Dimensions
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd65809f-b519-4dd9-aca7-bc452cb6adae · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values INSPECT: An Extensible Toolkit for AI Behavior Evaluation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c5cdae8-08c2-4c38-a138-b11566074059 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Governance in Agentic Workflows: Leveraging LLMs as Oversight Agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f34152c1-3bdc-4796-b437-1ecc6eaf0dc8 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c56c63fe-b2f0-40db-8b6f-319be09db8c4 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f226e5-346c-4313-bf58-63314a39d76c · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f47fd653-d640-4acc-b755-e93a1f270232 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Almost Surely Safe Alignment of Large Language Models at Inference-Time
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 535df10b-43a8-4d3e-afd6-35257bf5828b · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Dynamic Search for Inference-Time Alignment in Diffusion Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad083c49-e7a3-48d4-8c78-de4e9196656b · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f7c262-ade5-4959-a0a1-84e014de2c97 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Constitutional AI: Harmlessness from AI Feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae553b9-9542-4f09-986a-e2bb36766dc6 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0307194a-196c-4c9a-adca-c6c142cca4c6 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 171bc1d9-efa4-4d45-85cc-b610962d7f10 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values AI Control: Improving Safety Despite Intentional Subversion
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f222863a-22fb-447c-badb-4127e9498e77 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values OpenAI x DFT: The First Moral Graph
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ff4d278-26db-4e2f-a2be-693040d26acf · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Integrity
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27bab82c-a0dc-4c0c-ad58-3d049b54f333 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The Global Landscape of AI Ethics Guidelines
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d78eeaf-2263-4a86-9a54-c9b9b492120d · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values WhatsApp MCP Exploited: Exfiltrating Your Message History via MCP
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 48e2fdca-9791-46d1-a005-6bdfaca00cf9 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values MCP Security Notification: Tool Poisoning Attacks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20244967-148d-46a2-bb57-cdfbdc0be7ec · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31272682-141c-4702-a791-551e22302342 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Emergent Misalignment
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f40ad55-bb11-4b30-bc91-b86935cea64a · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Plurality
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 37621a40-2612-4ca3-812b-e77b1cdf5562 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Plurality: A Taxonomy for Pluralistic AI
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f6e36b4-7098-4f6c-9730-c2637ec23bd0 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc7ee6ef-2dca-4126-922c-1c188a5c2c2e · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values OASIS: Open Agent Social Interaction Simulations with One Million Agents
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b915a17-c0ea-45d4-93aa-e2fcfdfb4806 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Project Sid: Many-agent simulations toward AI civilization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106163d8-0eb8-450e-9b6f-081131884779 · outbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.