Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T15:26:23.290009Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.07918.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T15:26:23.290009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 821096e3-a873-4fe4-8478-c3123d7b18a2 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Barrick and Michael K
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a30391ca-3121-4b77-9a16-ddb254bec8d4 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Emergent Misalignment : Narrow finetuning can produce broadly misaligned LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 749af3db-bdfd-403d-bc3d-8ffabca8fdd1 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 20bea938-25bf-41aa-80d9-488db14c226a · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 25068ab6-9db4-44a4-a01e-1714c56d97b2 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 06da8fcd-fc02-4741-b04e-e45733d52699 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 240ccd28-d064-4a20-8c33-dbedea12db84 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Deep reinforcement learning from human preferences
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f8736ea1-80aa-42a2-b24a-850c7505c59c · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7dfadebb-76e8-40ed-b95f-16bce40796b8 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 49b741f9-7b42-4a6d-93d5-dd136354a5fa · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits du Plessis and Gideon P
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 80ae4665-bb36-466d-88e6-49e455d713df · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Psychometric personality shaping modulates capabilities and safety in language models, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4177629c-db75-4ca3-a86e-f165a8fabbe7 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation abe71d70-2a2f-4db4-8acd-caf8cca14887 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Goldberg, John A
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 43dc32dc-55e5-489c-be48-115524d23508 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Explaining and Harnessing Adversarial Examples
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 71ccc23e-eeee-43dc-b525-196efe517bf7 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Self-Assessment Tests are Unreliable Measures of LLM Personality
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 060a4676-9620-4f43-913c-6f21d404f175 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Measuring Massive Multitask Language Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b47e93bf-14c4-4c01-82d8-55f57aa76a87 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits International personality item pool: Alphabetical list of items
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0310fc50-cfa6-42f9-8f08-c536caf3c83e · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2a34e8c7-3430-419e-b6c3-9ef39e4abb87 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Evaluating and inducing personality in pre-trained language models, 2022
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 22a05933-eb2e-4f28-8965-ab42590079a0 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3137c19b-a412-43d6-bdeb-21ae7688e461 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c8a8c345-c191-44ab-beff-db1957af4265 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c767a605-2786-4fbf-bc53-ae930fd91568 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 98b70c22-b2b1-4f3e-bd68-3c973054c2f5 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits arXiv:2601.10387 [cs]
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 11766f72-aa03-40f3-a89b-79430d42efe4 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Stable and explainable personality trait evaluation in large language models with internal activations, 2026
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fe6c7a1a-57e5-4e9b-b146-7ec5e82f72fa · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 90b7e137-a4eb-4bad-978e-f8fd6343c4da · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 31fd26ec-c5a2-4b85-86c8-ce625b437b9a · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 73398e14-ad66-402f-8385-039860e29775 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4c5dd018-3fd8-4e9f-a1d1-6e7658c60902 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1954f2f3-4241-4142-a245-fcda157ab2f7 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 57f8316b-b701-45a6-8c74-ddc6f9191298 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6487dcc1-2ac4-4d81-8c13-60b2dc6d995a · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Red Teaming Language Models with Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4503ec75-b14e-4b7c-b05f-561353a21df0 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits tinyBenchmarks: evaluating LLMs with fewer examples
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 23b57278-cdf1-4e03-8c8a-03309603984e · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits InProceed- ings of the 32nd International Conference on Neu- ral Information Processing Systems, NIPS’18, page 5732–5741, Red Hook, NY , USA
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a07e89c1-7bdf-487e-92ba-ee7b96d14961 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Position: Adversarial ML for LLMs Is Not Making Any Progress
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 022ce466-ebc4-422d-833a-d5511697f8f2 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0a34a515-b193-45e0-8450-b95ee5182d2e · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e23e69b8-bef5-44b2-a51a-8e2cd5902916 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits A coin flip for safety: Llm judges fail to reliably measure adversarial robustness
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e8aa52b9-50b7-4857-ba42-2e5c7aaf4568 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9bf7fdef-e646-46e4-a844-e99d675fb663 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits A psychometric framework for evaluating and shaping personality traits in large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e055885f-daed-490f-bbe7-51d7258e870d · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e68132a7-490c-4fc8-a276-fd4a3bb91eb2 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c85018f4-f52f-471a-92eb-af959a224b44 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Simms, Lewis R
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b356fd59-70a3-45cb-b68a-317b411eb8a3 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Evaluating alignment of behavioral dispositions in llms
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c35de6d0-c28d-49a9-a122-40e5fe3aeb44 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 38d1709e-1c21-400a-8043-5f78fb9bbfdb · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits arXiv preprint arXiv:2506.19823 , year =
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6a0614ed-c35e-4664-8a47-c51f60777f18 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6cb77377-5147-4dc2-8859-5a74d98068fb · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Efficient Adversarial Training in LLMs with Continuous Attacks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4fcc0c92-4f51-4d48-8286-1232dd0c9bf7 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Bullying the Machine: How Personas Increase LLM Vulnerability
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5082985f-e678-4227-bf83-8b5b129af0f0 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Qwen3 Technical Report
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ee1a0b98-dd8a-45a9-b9d0-7bf12bcd306d · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 497d297c-1db3-4b2c-9383-ca0925800bd5 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Robust LLM safeguarding via refusal feature adversarial training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b237e65-4076-4ba5-8819-f8306f54ed21 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Coifman, and Ruoming Jin
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0e0a87bf-c2cc-485f-97a2-956954b61dcd · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Personality Alignment of Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b64926a1-9fdf-47c1-a36b-38a2ef3998b2 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3edfe9bd-0143-4ded-bd56-737dec0ccc61 · outbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Representation Engineering: A Top-Down Approach to AI Transparency
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.