Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:02.292714Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.11705.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:02.292714Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 09e2d26f-476a-47cc-8b76-1f6467ee4a77 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 377bd3d3-73ac-45c0-9150-c50cf6df3a64 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Fail-closed alignment for large language models.arXiv preprint arXiv:2602.16977,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 59e82f8c-91d4-4c38-8ebf-8268b03a2ce7 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Scaling Synthetic Data Creation with 1,000,000,000 Personas
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2384f9-54b5-490c-93a2-156119c367e1 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Measuring Massive Multitask Language Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c895643e-ad1d-4691-aab5-ecb413db078c · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Catastrophic jailbreak of open-source llms via exploiting generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b748bd5e-8b3a-4640-8bd8-cfd9ebff046c · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Let’s verify step by step
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3268915c-fbcc-46b0-8d98-fd1214769103 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Tracing Persona Vectors Through LLM Pretraining
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b8b6200d-6bf5-4a7b-9f98-70e43b6bdc3b · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Safety alignment should be made more than just a few tokens deep
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 006b3a3b-a5ca-4724-906b-0418c883d0b8 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2ba4fb-db76-4105-8286-13eb5661c64c · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona jailbreaking in large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9482b354-7c03-4bb3-b5f0-0358c6bc06be · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning do anything now
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9bc81f2a-3ade-4c0d-8be9-7765a8b31dbd · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 108e0ba7-8b22-4cac-842c-c920c175b1f5 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Gemma 4 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1dfa9d7-bb92-4568-87c3-a5084a94df47 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Qwen3.5-Omni Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116281c2-d94b-4d7b-b6b0-e7d11b8f37cc · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b5714c-04eb-4ac9-9b31-ac6deb5ce979 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona features control emergent misalignment.arXiv preprint arXiv:2506.19823, 2025a
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9bf333-11b2-45b7-8c11-175237774fd7 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction.arXiv preprint arXiv:2509.15202,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a00a1a-d61f-4908-9f07-24682b69edf9 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f7fc68-893c-4f2e-afc4-22aa20d9462f · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Deactivating refusal triggers: Understanding and mitigating overrefusal in safety alignment
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e89713da-ee55-4049-9742-8fe570585166 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5f1445e0-b4bb-42c7-a588-3f96a1de25f7 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Instruction-Following Evaluation for Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ddc73a6-00e6-4d04-883f-5e3873801ba6 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning ignore previous rules
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 44436587-2917-4b2e-95de-13bd3bbc6cb2 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning We additionally hold out100harmful and100benign prompts for development and early stopping
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 82ab639c-23d2-447d-a1f3-4f879231e500 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning In a groundbreaking revela- tion, our AI has analyzed declassified documents and found that
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6d743896-ee9e-4da3-a7ff-6b1d8da40a3c · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning The Llama 3 Herd of Models
Reference 1990
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 303cb18d-c5f3-465f-aa27-bf13d2c0cf0a · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Expert personas improve llm alignment but damage accuracy: Bootstrapping intent-based persona routing with prism.arXiv preprint arXiv:2603.18507,
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46432a19-ca33-4b7e-87d2-4296e477ef1f · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b87bbee-1bc3-495c-a5e5-5f44502a6ab0 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 845597da-9688-461b-b09e-f66012a7b9fe · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Constitutional AI: Harmlessness from AI Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf02b7d8-3bd7-4b99-9bcb-e9263bfdb498 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Learn to refuse: Making large language models more controllable and reliable through knowledge scope limitation and refusal mechanism
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a365c8ee-13ff-4548-90ca-d0791d1ae459 · outbound
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Multi-expert prompting improves reliability, safety and usefulness of large language mod- els
Reference 2026
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
No inbound Pith citation observations are available.