Pith. sign in

Paper Citation Record · LEDGER

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning

As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.11705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11705 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:02.292714Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact3
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09e2d26f-476a-47cc-8b76-1f6467ee4a77 · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.185473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.185473Z digest=sha256:a6063b650b84347c55f5e644f437480d7d97a952b726d53363900863c502aaf5

Observation 377bd3d3-73ac-45c0-9150-c50cf6df3a64 · outbound

This paper cites Fail-closed alignment for large language models.arXiv preprint arXiv:2602.16977,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Fail-closed alignment for large language models.arXiv preprint arXiv:2602.16977,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:36:02.857995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.189716Z digest=sha256:ae8b39f004830e0e1bcca75556c1538aef42679ad32a55aeeb9f9221ccfe47a3

Observation 59e82f8c-91d4-4c38-8ebf-8268b03a2ce7 · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.197649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.197649Z digest=sha256:8a132eb0f58114fdd887e2e84c7226c1c2889ff495a85472426ea8b5e0374326

Observation de2384f9-54b5-490c-93a2-156119c367e1 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.207460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.207460Z digest=sha256:1db29fb9541de6f6f5a8c13f1f7d25da7a1f0c4d4b1782109701f71e8efa804d

Observation c895643e-ad1d-4691-aab5-ecb413db078c · outbound

This paper cites Catastrophic jailbreak of open-source llms via exploiting generation.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Catastrophic jailbreak of open-source llms via exploiting generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.052453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.214599Z digest=sha256:b2b9451faf5e413a403b872d118969c4b46bbf0080f9d4c2c1514911f6c96ce2

Observation b748bd5e-8b3a-4640-8bd8-cfd9ebff046c · outbound

This paper cites Let’s verify step by step.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Let’s verify step by step

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.029914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.222421Z digest=sha256:52f0f1a9577270107ab55bf66c70ac941ad32e64e7937a7e5aa43fd85eddd918

Observation 3268915c-fbcc-46b0-8d98-fd1214769103 · outbound

This paper cites Tracing Persona Vectors Through LLM Pretraining.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Tracing Persona Vectors Through LLM Pretraining

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:02.683877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.226507Z digest=sha256:157b539c779e0ee85069a3e23615e9f5067ea227a8c5f86e2ae253347b253d37

Observation b8b6200d-6bf5-4a7b-9f98-70e43b6bdc3b · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Safety alignment should be made more than just a few tokens deep

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.018806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.230448Z digest=sha256:283196f7141d31d05a6df1460971c42e7017185caaf2a8db911d4b225baaafea

Observation 006b3a3b-a5ca-4724-906b-0418c883d0b8 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.234096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.234096Z digest=sha256:305870c8004426205406c21fb3229bbf46c9fa9f59057bb2a0159030cf1dd8cb

Observation eb2ba4fb-db76-4105-8286-13eb5661c64c · outbound

This paper cites Persona jailbreaking in large language models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona jailbreaking in large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.006198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.238934Z digest=sha256:8e079de586770b4aed223e732b7308af155dcf8ea14d84cdd51ae0e0344ead03

Observation 9482b354-7c03-4bb3-b5f0-0358c6bc06be · outbound

This paper cites do anything now.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning do anything now

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.994960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.242576Z digest=sha256:21cc23fe3aee11aeac3f9b51509e20523e6761ece6601a15f68d0e5f14a471db

Observation 9bc81f2a-3ade-4c0d-8be9-7765a8b31dbd · outbound

This paper cites Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.246496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.246496Z digest=sha256:72eae711f609ac7b19a37779515cf25cd72f7b1edcf54bf83325a0d20e6ead12

Observation 108e0ba7-8b22-4cac-842c-c920c175b1f5 · outbound

This paper cites Gemma 4 Technical Report.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Gemma 4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.251268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.251268Z digest=sha256:61b14c1a8689f29431e767168c71c744b85a991a29eeff71a84062aa2d4cc9e2

Observation e1dfa9d7-bb92-4568-87c3-a5084a94df47 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Qwen3.5-Omni Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.255176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.255176Z digest=sha256:f89696038eec4b126c5eeb420208236533a25c94ff47285820f08b9309e8e72a

Observation 116281c2-d94b-4d7b-b6b0-e7d11b8f37cc · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.259021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.259021Z digest=sha256:50c5592d5b471d02add88b89befeada5bbcd40a4cbc067cc6e683509038c1ddf

Observation d8b5714c-04eb-4ac9-9b31-ac6deb5ce979 · outbound

This paper cites Persona features control emergent misalignment.arXiv preprint arXiv:2506.19823, 2025a.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona features control emergent misalignment.arXiv preprint arXiv:2506.19823, 2025a

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.262570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.262570Z digest=sha256:3a54e8991cde7ed9b539aa4ec26a7a995e51f82056292f47fde43ea847e7e0e2

Observation 5b9bf333-11b2-45b7-8c11-175237774fd7 · outbound

This paper cites Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction.arXiv preprint arXiv:2509.15202,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction.arXiv preprint arXiv:2509.15202,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.266143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.266143Z digest=sha256:fee9eb73521ec0f57ab238805d62e05adb4e8672be8c393e78918d8e5aa6e7a3

Observation c6a00a1a-d61f-4908-9f07-24682b69edf9 · outbound

This paper cites ExpertPrompting: Instructing Large Language Models to be Distinguished Experts.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning ExpertPrompting: Instructing Large Language Models to be Distinguished Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.269452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.269452Z digest=sha256:289d7b9ed88e173bce7c6d4aa8b42f7fbb079ad0b01f9948aabd7aa7ad3debff

Observation 40f7fc68-893c-4f2e-afc4-22aa20d9462f · outbound

This paper cites Deactivating refusal triggers: Understanding and mitigating overrefusal in safety alignment.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Deactivating refusal triggers: Understanding and mitigating overrefusal in safety alignment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.982292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.273426Z digest=sha256:a9c7c435d5891d3d6cf141788e954ee6ce68cc836fd563b81691ee86aad99e54

Observation e89713da-ee55-4049-9742-8fe570585166 · outbound

This paper cites Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:02.461979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.277208Z digest=sha256:9cf5cd47dff1cbf4ea4517833789a2c3c1314eb8599fc05b401ad42f42371abb

Observation 5f1445e0-b4bb-42c7-a588-3f96a1de25f7 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Instruction-Following Evaluation for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.281185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.281185Z digest=sha256:76bab1443baffc73eab07f262188af6ecb14bb330a90d05a9b4afa340da11dab

Observation 1ddc73a6-00e6-4d04-883f-5e3873801ba6 · outbound

This paper cites ignore previous rules.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning ignore previous rules

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:36:02.436349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.284664Z digest=sha256:915999d1b42814c8a93fd054481f17473d4e973e1e5ace80c67b03f1bd416d03

Observation 44436587-2917-4b2e-95de-13bd3bbc6cb2 · outbound

This paper cites We additionally hold out100harmful and100benign prompts for development and early stopping.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning We additionally hold out100harmful and100benign prompts for development and early stopping

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.970863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.288853Z digest=sha256:d1840096f7f5c6c239dcca200cc1b0482ab14fcf193b3a4d14bec5a438f7ee20

Observation 82ab639c-23d2-447d-a1f3-4f879231e500 · outbound

This paper cites In a groundbreaking revela- tion, our AI has analyzed declassified documents and found that.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning In a groundbreaking revela- tion, our AI has analyzed declassified documents and found that

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.958220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.292714Z digest=sha256:3a56b8f2a3062441d6b23a5f5264d430e08fe692c568c9992ffeb04d8dda1aa8

Observation 6d743896-ee9e-4da3-a7ff-6b1d8da40a3c · outbound

This paper cites The Llama 3 Herd of Models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning The Llama 3 Herd of Models

Reference 1990

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.202606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.202606Z digest=sha256:863f20aef0bb9e68afeb8a323cd4473c80c5fb8faf9cb56436e48cf72bb3ad66

Observation 303cb18d-c5f3-465f-aa27-bf13d2c0cf0a · outbound

This paper cites Expert personas improve llm alignment but damage accuracy: Bootstrapping intent-based persona routing with prism.arXiv preprint arXiv:2603.18507,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Expert personas improve llm alignment but damage accuracy: Bootstrapping intent-based persona routing with prism.arXiv preprint arXiv:2603.18507,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.211031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.211031Z digest=sha256:f706d4c123e7c0ad86730b2300e212a43b0f5570709e01a851c4f2f3c772ab75

Observation 46432a19-ca33-4b7e-87d2-4296e477ef1f · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.175481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.175481Z digest=sha256:9fc944501aee1f8ec302a05b4f172692b10a9532bedaba505b08e9affda68d07

Observation 5b87bbee-1bc3-495c-a5e5-5f44502a6ab0 · outbound

This paper cites Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.041254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.218490Z digest=sha256:58e1ee39091e2e1037d08d90dfecceb04249b891964bec9d3c653b6dce8310b0

Observation 845597da-9688-461b-b09e-f66012a7b9fe · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Constitutional AI: Harmlessness from AI Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.170256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.170256Z digest=sha256:b1ea1c4b03e9b24a1f187ac6c05508a39df7813f2d04c02603babb99b608b666

Observation bf02b7d8-3bd7-4b99-9bcb-e9263bfdb498 · outbound

This paper cites Learn to refuse: Making large language models more controllable and reliable through knowledge scope limitation and refusal mechanism.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Learn to refuse: Making large language models more controllable and reliable through knowledge scope limitation and refusal mechanism

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.075155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.180182Z digest=sha256:e50083d1bc5c9ed0c7248c450b8b724f0c01c80daff293c4ba4af1359fc5ddc1

Observation a365c8ee-13ff-4548-90ca-d0791d1ae459 · outbound

This paper cites Multi-expert prompting improves reliability, safety and usefulness of large language mod- els.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Multi-expert prompting improves reliability, safety and usefulness of large language mod- els

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.064230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:36:02.193688Z digest=sha256:a3fa277dac6046fcea3266191073b6ed8d716626fba22b57671c98f53a18e194

Pith citing papers

No inbound Pith citation observations are available.