Pith. sign in

Paper Citation Record · LEDGER

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2506.13774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13774 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:07.337382Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03ffc49a-1a7b-429c-8020-54a96502feb0 · outbound

This paper cites Artificial Intelligence, Values, and Alignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Artificial Intelligence, Values, and Alignment

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.261527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.102830Z digest=sha256:e6f7401c0b917f13c92944ddbaf7f62783887ccf8ec17884780da091d32d48b6

Observation 93026cbf-d874-413a-9bd0-b2651e5c8a3b · outbound

This paper cites Artificial Morality: Top -down, Bottom-up, and Hybrid Approaches.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Artificial Morality: Top -down, Bottom-up, and Hybrid Approaches

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.247705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.113520Z digest=sha256:9bed11efb4ec4c91cd7d4f17e98e8bea0d6ce6d9c61544a656f41c3b9823fbdb

Observation e4ec19f7-68a5-4123-90ff-cb5f9267b597 · outbound

This paper cites Translating Principles into Practices of Digital Ethics: Five Risks of Being Unethical.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Translating Principles into Practices of Digital Ethics: Five Risks of Being Unethical

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.232427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.119100Z digest=sha256:d46741bbcc05a1ab157ff95665fa077c1856683a64174e3141f793c8dd02ea50

Observation cab35d18-a447-4352-ac03-fd37e5153ac7 · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.123823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.123823Z digest=sha256:950d97bac0dadd549e557d551d14502657cd3189fbe16a5265314d0e163e85f0

Observation c5354f90-fad6-4bb2-96cf-a66ca53ebf2d · outbound

This paper cites Personalized Large Language Models.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Personalized Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.128902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.128902Z digest=sha256:ee97c7148beb417d6182afa7c8df21129ed252216db6ebfa6cb351a3b9a78236

Observation 49e8007a-9ada-423f-8dbc-f1bc1335f848 · outbound

This paper cites Towards an End -to-End Personal Fine -Tuning Framework for AI Value Alignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Towards an End -to-End Personal Fine -Tuning Framework for AI Value Alignment

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.218849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.133884Z digest=sha256:ec79bcf3d95e86e54976a67ee17354702ec9cb276cbaa0304ef2379032e49821

Observation a1f78d6b-2a43-4697-b583-f52683d34c35 · outbound

This paper cites Safer Agentic AI.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Safer Agentic AI

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.204958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.139059Z digest=sha256:4169ceac4cdcb1c480617adce39de2e47205d4fefcbc8d1e6dbfa70d3c0afac6

Observation ab4fc287-4bcb-4053-b2df-466b9746a33d · outbound

This paper cites Introducing the Model Context Protocol.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Introducing the Model Context Protocol

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.190940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.143540Z digest=sha256:771ccf3ec5211124d0c7997f0ea37fb41d951941f3750d861e0fd20b73a25675

Observation 0b0b6157-c207-4c1b-b451-66d573f3884b · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Deep Reinforcement Learning from Human Preferences

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.176703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.147943Z digest=sha256:1963fecc4be4df99ec0495ce4b821dee7568552e89e13672384cfc3cf9b8c9a4

Observation c1a6a662-4abe-47f4-b593-49c0d6ede551 · outbound

This paper cites Confabulation: The Surprising Value of Large Language Model Hallucinations.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Confabulation: The Surprising Value of Large Language Model Hallucinations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.152276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.152276Z digest=sha256:44586319aed738e3c2d67d7352417949b065428ebe1a51aa5e4d337549a02011

Observation 75f881bc-d2a2-4cc7-a983-2f2387095612 · outbound

This paper cites Choice Vectors: Streamlining Personal AI Alignment Through Binary Selection.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Choice Vectors: Streamlining Personal AI Alignment Through Binary Selection

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.161821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.157090Z digest=sha256:98fef6c41b0a0258efdfc65405991546592f74bba577bc227f8660769f59a65a

Observation f8843945-5b73-4fb6-8fdd-6eacf9e521a7 · outbound

This paper cites an unresolved cited work.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:43:08.147667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.161465Z digest=sha256:3c6713d3324e8ac101b0c45584d3de4a191c8989db4d44632f82b7ab807ffeca

Observation 232768fa-1ba1-48b4-b37f-24a77dad24c0 · outbound

This paper cites Universality of Representation in Biologic al and Artificial Neural Networks.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Universality of Representation in Biologic al and Artificial Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.165649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.165649Z digest=sha256:0531853e7e02bc5a824e76a9564aa36f569900503b2a440ef0ab395031bdf310

Observation b5015ad3-26fb-48f4-8a6f-ff1489e4c7a9 · outbound

This paper cites The neural bases of cognitive conflict and control in mo ral judgment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The neural bases of cognitive conflict and control in mo ral judgment

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.133544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.169921Z digest=sha256:ab755b1038726ed779c726f533aa90d891da2e26201964582c558945ac9f37be

Observation 5f808b93-0ce8-402b-a21e-fc5f0480cb73 · outbound

This paper cites The neural basis of human social values: Evidence from functional MRI.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The neural basis of human social values: Evidence from functional MRI

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.118460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.174334Z digest=sha256:1a4be3402b2e6c13b399cc36cb3b53471ee68bee2dc942641efb188029e73b0f

Observation 2d831fb6-b9b9-40aa-af80-9648e02cea1c · outbound

This paper cites A Cognitive Theory of Consciousness: The Workspace of the Mind; Cambridge University Press: Cambridge, UK, 1988.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values A Cognitive Theory of Consciousness: The Workspace of the Mind; Cambridge University Press: Cambridge, UK, 1988

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.103481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.179006Z digest=sha256:2ba499fb0d949e0a3f628ad85ab85510be95ab2bcc9e4820a1fc66cc8d3f72e9

Observation 6886e3ee-5210-46b2-b28e-efef9e830027 · outbound

This paper cites Unified Theories of Cognition; Harvard University Press: Cambridge, MA, USA, 1990.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unified Theories of Cognition; Harvard University Press: Cambridge, MA, USA, 1990

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.088209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.183568Z digest=sha256:ee3bdc0a1ede23c46702093d6db6c0d9048abc928251d97f5eb8581c53d505ab

Observation bec20e91-f154-4787-a276-9d29e7bc19d6 · outbound

This paper cites Revealing economic facts: LLMs know more than they say.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Revealing economic facts: LLMs know more than they say

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:43:07.780589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.188289Z digest=sha256:ab2f04188ea59cc76c6e6e71bfc80f646beb02e233cee5e1a04dd7af8d7ec059

Observation 565c71ff-9edc-481f-a63f-4b03b833470d · outbound

This paper cites ShieldGemma 2: Robust and Tractable Image Content Moderation.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values ShieldGemma 2: Robust and Tractable Image Content Moderation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.193249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.193249Z digest=sha256:d7600a6311d9d121542732c3330c9ca35484f30e0de974763bb563461922d727

Observation 40c700b9-bd04-4265-a0d1-417a7d2802c3 · outbound

This paper cites Superego -Agent LGDemo (Branch: Fastapi_Mcp).

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Superego -Agent LGDemo (Branch: Fastapi_Mcp)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.073012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.198921Z digest=sha256:6c4b8eb44ddc6d76cec4b7959fd7226d5a44fc60ccfb39b1ebffa50186726d5c

Observation c55222a1-896d-459a-9f3e-dfd48bfa9033 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.203855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.203855Z digest=sha256:bd29590b4448291174740f99a91a64a02e552af510de36c4ec9adeaa11a4f9fa

Observation 218dadac-60a5-4f85-909e-cf5e6a25d19e · outbound

This paper cites AgentHarm: A benchmark for measuring harmfulness of LLM agents.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values AgentHarm: A benchmark for measuring harmfulness of LLM agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.058305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.209304Z digest=sha256:27c85c39ef24195cbf32c759f39db7f2652c5e6104675b128eba247c97953be4

Observation 0d21dcae-dcab-4951-b7a8-6e0cf673c985 · outbound

This paper cites Do the rewards justify the means? Measuring trade -offs between rewards and ethical behavior in the Machiavelli benchmark.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Do the rewards justify the means? Measuring trade -offs between rewards and ethical behavior in the Machiavelli benchmark

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.043769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.213631Z digest=sha256:fefd4df8634c623804b8770e600383744c8a5983d89a695d94111ac709af01dd

Observation 7d230248-3583-4771-9434-10d2f6bcc2e5 · outbound

This paper cites Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.218428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.218428Z digest=sha256:b88fa9e33c5407ffd5d1b0b4f67b50b4f3e0d63a7b45770490c6619b5fe9a975

Observation b9a1bd80-b170-4952-ac58-bb35175f738c · outbound

This paper cites Vijil Test Library: Evaluating LLM Trustworthiness Across Eight Dimensions.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Vijil Test Library: Evaluating LLM Trustworthiness Across Eight Dimensions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.029110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.223161Z digest=sha256:370f3f7a295779e98e102d79483d2bc15354815ea51e8513877ed9573bfa53b9

Observation dd65809f-b519-4dd9-aca7-bc452cb6adae · outbound

This paper cites INSPECT: An Extensible Toolkit for AI Behavior Evaluation.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values INSPECT: An Extensible Toolkit for AI Behavior Evaluation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:08.013655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.227370Z digest=sha256:bafebee85b6a967ce2c720e781b900b4ce75354963d646a30c98b5d41c1bbfdd

Observation 5c5cdae8-08c2-4c38-a138-b11566074059 · outbound

This paper cites Governance in Agentic Workflows: Leveraging LLMs as Oversight Agents.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Governance in Agentic Workflows: Leveraging LLMs as Oversight Agents

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.999467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.231834Z digest=sha256:0cad9894e395512f72aef27efaddd36b20f30eb13683b41992c7f906dd854bd1

Observation f34152c1-3bdc-4796-b437-1ecc6eaf0dc8 · outbound

This paper cites Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.235972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.235972Z digest=sha256:8d3822c9efdea63377dacae0bc7f3a74a96951813dcfc095d7a34176c1a59d53

Observation c56c63fe-b2f0-40db-8b6f-319be09db8c4 · outbound

This paper cites Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.240615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.240615Z digest=sha256:9565cc2e445757f8b6cf3e4922d3709adc99397c5c9d04219cfbb922614f3237

Observation 01f226e5-346c-4313-bf58-63314a39d76c · outbound

This paper cites InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.245116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.245116Z digest=sha256:83eccd0dd4effe37bf7ad73f3cb2e9a95149ddc71991a118098308be005e8312

Observation f47fd653-d640-4acc-b755-e93a1f270232 · outbound

This paper cites On Almost Surely Safe Alignment of Large Language Models at Inference-Time.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Almost Surely Safe Alignment of Large Language Models at Inference-Time

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:43:07.630186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.249475Z digest=sha256:5bf910618f160910fdac07653f6bda2024cf652dd046eb1bfd6aa7536637d270

Observation 535df10b-43a8-4d3e-afd6-35257bf5828b · outbound

This paper cites Dynamic Search for Inference-Time Alignment in Diffusion Models.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Dynamic Search for Inference-Time Alignment in Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.253673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.253673Z digest=sha256:f3c79977d0e3c6fa2f6c044bc8646d18fd401a5a608ff0a76934291bf2cc379d

Observation ad083c49-e7a3-48d4-8c78-de4e9196656b · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.258162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.258162Z digest=sha256:8d41976cb96366aaa68515352150b0fdcf529b7359dbca69480315517d14682f

Observation b7f7c262-ade5-4959-a0a1-84e014de2c97 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Constitutional AI: Harmlessness from AI Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.262772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.262772Z digest=sha256:d53878783b789272ea848c6b9cb2d6a0d71d4541fbe37fffc1022904fb277279

Observation 7ae553b9-9542-4f09-986a-e2bb36766dc6 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.267381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.267381Z digest=sha256:7c33cb19d910f2d71536427b0fdd171d8372a13a5851e36128e5801ffbf97e11

Observation 0307194a-196c-4c9a-adca-c6c142cca4c6 · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.984868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.272296Z digest=sha256:4de4533ea5cf00027ef6650d9eb93460895b76542933300b2392d52e7e122bdc

Observation 171bc1d9-efa4-4d45-85cc-b610962d7f10 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values AI Control: Improving Safety Despite Intentional Subversion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.278246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.278246Z digest=sha256:22f4a0e296364097196bf338fc069b38be3b02887efb2042cc2585020d467655

Observation f222863a-22fb-447c-badb-4127e9498e77 · outbound

This paper cites OpenAI x DFT: The First Moral Graph.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values OpenAI x DFT: The First Moral Graph

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.970374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.283191Z digest=sha256:7a8f4e686fe85356972f9b596145a44cfe6131af3e4d298a9198d3a994deff17

Observation 8ff4d278-26db-4e2f-a2be-693040d26acf · outbound

This paper cites Model Integrity.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Integrity

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.956049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.287785Z digest=sha256:bfa919bd42a21d6dde17817ceaf5ef5416c5c48ba2d0b9f56f45e37da11c6787

Observation 27bab82c-a0dc-4c0c-ad58-3d049b54f333 · outbound

This paper cites The Global Landscape of AI Ethics Guidelines.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values The Global Landscape of AI Ethics Guidelines

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.941746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.292294Z digest=sha256:191f1b1beb34f77d843cc1a338de1ee23190756475f5ad0bb401dfa119b93f16

Observation 5d78eeaf-2263-4a86-9a54-c9b9b492120d · outbound

This paper cites WhatsApp MCP Exploited: Exfiltrating Your Message History via MCP.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values WhatsApp MCP Exploited: Exfiltrating Your Message History via MCP

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.926146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.296584Z digest=sha256:711a2844b438d59c5f19286063cd383931c97c10ac380437a8e1c4248e87fc20

Observation 48e2fdca-9791-46d1-a005-6bdfaca00cf9 · outbound

This paper cites MCP Security Notification: Tool Poisoning Attacks.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values MCP Security Notification: Tool Poisoning Attacks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.895246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.306568Z digest=sha256:6f6407760dc0861343466a63a0fe9248e42866dacc0f6d5a431eace0d3293655

Observation 20244967-148d-46a2-bb57-cdfbdc0be7ec · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.311471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.311471Z digest=sha256:89e21ca2d7f94cddefe6ed34d6be80f213a6bb9753d2d643d3d94cdb4a69ebb7

Observation 31272682-141c-4702-a791-551e22302342 · outbound

This paper cites On Emergent Misalignment.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Emergent Misalignment

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.880226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.315617Z digest=sha256:0af311189b57859f53206a0b3182616830930d530e32c9d758f5ad9340f98124

Observation 5f40ad55-bb11-4b30-bc91-b86935cea64a · outbound

This paper cites Model Plurality.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Plurality

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.864809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.319836Z digest=sha256:d3755a5fa13c313c95624ab994bd467e7b35917f23cb12c75978494ea06e48b4

Observation 37621a40-2612-4ca3-812b-e77b1cdf5562 · outbound

This paper cites Model Plurality: A Taxonomy for Pluralistic AI.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Model Plurality: A Taxonomy for Pluralistic AI

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:07.849749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.324191Z digest=sha256:c48019f8225a80d38fdfc15a15dde578c54f69671574a1e1efdc2a465af7304e

Observation 6f6e36b4-7098-4f6c-9730-c2637ec23bd0 · outbound

This paper cites Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.328340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.328340Z digest=sha256:10793edbc3226c0fff9d4331cc1d79ed9bf7ec9212a2f0e74db9f76b3815176f

Observation dc7ee6ef-2dca-4126-922c-1c188a5c2c2e · outbound

This paper cites OASIS: Open Agent Social Interaction Simulations with One Million Agents.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values OASIS: Open Agent Social Interaction Simulations with One Million Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.332394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.332394Z digest=sha256:69ca7686fae471dab3cffc8fd291b10a1060ae3535f509f6911ba13aee391b6a

Observation 1b915a17-c0ea-45d4-93aa-e2fcfdfb4806 · outbound

This paper cites Project Sid: Many-agent simulations toward AI civilization.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Project Sid: Many-agent simulations toward AI civilization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.337382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.337382Z digest=sha256:a49e6eeea1370a2374fa361aca712735fdc26b7d93f178de9cdba358c4466699

Observation 106163d8-0eb8-450e-9b6f-081131884779 · outbound

This paper cites an unresolved cited work.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:43:07.910198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:43:07.301344Z digest=sha256:3766f5b843fa10fc57a8a32ed9a52ac157960e7d942bf33d45d7bd0f1c2a8d76

Pith citing papers

No inbound Pith citation observations are available.