Pith. sign in

Paper Citation Record · LEDGER

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2502.02153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02153 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:14:34.157748Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e99a51a-78c5-41c9-b065-1cd0e964302c · outbound

This paper cites GPT-4 Technical Report.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:33.962283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:33.962283Z digest=sha256:26407e9bcd92ab7789eec41cb45704ebf81f93bdf8a14dc2fe7ad4819f75e669

Observation 46e8b5c1-bf95-4261-af75-94fec11633a3 · outbound

This paper cites Concrete Problems in AI Safety.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Concrete Problems in AI Safety

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:33.966846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:33.966846Z digest=sha256:8f94834edf0206a76eb8e95539414c5a41683cd292f623c19e11ee64a5869977

Observation e9ec5777-ad42-4e52-a73e-8723cdbbbe3c · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing A general theoretical paradigm to understand learning from human preferences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.796288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:33.971023Z digest=sha256:9fec35c44a5cebd2d72d7a157de18ff38f3f02b97f03ff1f210262e35754abd0

Observation 5c7e0272-d3be-44bc-bfff-8c652f2e3c3f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:33.974327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:33.974327Z digest=sha256:d56650062cc5b7c92e36ebe2e5ef6ef2e67e1d17261ded31f2dc6efae0e0566b

Observation 184e74a8-9b08-463c-850f-5e5d9544af9b · outbound

This paper cites The ethics of artificial intelligence.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing The ethics of artificial intelligence

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.786736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:33.977968Z digest=sha256:3f8e6caa01c0b1677d5967ec8332bc8d7fff215c952c74db8dbb97a380cf04fa

Observation 7db00aeb-e6a8-4640-8b25-34ecbf99858e · outbound

This paper cites Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:33.981567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:33.981567Z digest=sha256:31c518a3ea0b10ce80b09cb5bbefaf5c41c1d33b1fd31863a083e7ddc6f29cf1

Observation e750a8bd-f1da-499c-9799-0934458e5bd8 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:33.985207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:33.985207Z digest=sha256:89820bcd58238c61ed65bb603db15964a91bd2cb05c0ed34c01b532f529bd26e

Observation 11c5e9c0-6400-4e2e-94bb-ab5198eadf68 · outbound

This paper cites Deep reinforcement learning from human preferences.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Deep reinforcement learning from human preferences

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.777594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:33.988726Z digest=sha256:ef14375d7df565973cc6847c0418ded779fd31cbbf8771b800251fe14ff66562

Observation fc7cf301-b09c-47d7-b707-001fba0ffc63 · outbound

This paper cites Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:33.992120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:33.992120Z digest=sha256:8b7e81287fc6f170948350ba0cb30a8acda05e6f0273cda702c3107491a6fe67

Observation dfab5b30-930c-4324-b9f7-71afa4d2da0a · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Safe RLHF : Safe reinforcement learning from human feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.767615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:33.997032Z digest=sha256:4ef1880b59058a8bb7cae9c79c6614ed7924614883ef9cecf8272fd271de0b4a

Observation cb5db034-2f19-4d33-9cc9-24dbfe1d4142 · outbound

This paper cites Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:14:34.570965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.000270Z digest=sha256:38cd5979d35c7c1c271aecd7a0ee49007fda0cdb397148d7099cb42b1d482713

Observation b13d94c2-4825-4279-b487-e61718bb53ce · outbound

This paper cites The Llama 3 Herd of Models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.003785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.003785Z digest=sha256:537aa1113568718fd535170f35fa9df42e074ab19c56c06b62bbd5d0ff342b4d

Observation 6e1a8042-a0c2-4117-852c-1aca7cb31e64 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing KTO: Model Alignment as Prospect Theoretic Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.007032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.007032Z digest=sha256:277866674df46f1ca4db855eb87f6789b4e577f70181483d0d29bfc166e9f224

Observation 9f96e0cd-a1b2-423d-9820-ca7046ffb9a6 · outbound

This paper cites Pal: Program-aided language models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Pal: Program-aided language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.758369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.010357Z digest=sha256:e7cce61cb7f1ec3a267c64f81238a53ea04ed5b6e5401fa56522626f1085eee1

Observation 9101effa-532c-40d7-b501-1987bd7cde2e · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.013524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.013524Z digest=sha256:4521a66203d84b95c97cccb86bd606a1ede375c20a18305df125ce7f67203847

Observation af0096dc-7c00-40b2-abc7-e90bd0b509c7 · outbound

This paper cites Measuring massive multitask language understanding.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Measuring massive multitask language understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.016884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.016884Z digest=sha256:d8c1c1f2e2960b061625ac31d7dae4dfa06ed1cef5d8c4659f049e0c022cce37

Observation 902cc25e-67ee-4d32-b70d-0278c7fc7eeb · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing ORPO: Monolithic Preference Optimization without Reference Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.020591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.020591Z digest=sha256:16ea7eb58f79846bddc56f1396e6888920e4e02ed7bb5763b82adb0b8c55c759

Observation fdb16407-657e-4b5d-b21f-9dcc3433fc89 · outbound

This paper cites One-Shot Safety Alignment for Large Language Models via Optimal Dualization.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing One-Shot Safety Alignment for Large Language Models via Optimal Dualization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.024260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.024260Z digest=sha256:1fca962448e59bd607f7e01cd1740f2952547b736fb31149120f7756e18b58e1

Observation 98c68720-f744-4ff2-869d-311c1b393377 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.027501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.027501Z digest=sha256:9edbb13b52794838a6f997b36448b50df3c016f7d3e25f1ce69c29f40f382dfa

Observation 3c1d3a8a-76af-43e3-993a-6aaa39a579f6 · outbound

This paper cites Reference point specification in hypervolume calculation for fair comparison and efficient search.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Reference point specification in hypervolume calculation for fair comparison and efficient search

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.743349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.031002Z digest=sha256:71c8434ea2e21125f50035adb7d7eb853e1a86677b0fe88b1d9572a9c8a34d84

Observation 33c186d5-3f4b-4479-ad01-e69a2adf1ffd · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing AI Alignment: A Comprehensive Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.034236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.034236Z digest=sha256:e2830ebfe01f56f02b2ef0571a4fa31a3a0a594c9e3166820d4972942d29fdc1

Observation 44ab9950-d943-44e9-9e48-8df82a5721f4 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.037586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.037586Z digest=sha256:fc7fd31476a4d32fc62cef013736a0e7aa6f294ef14b2dc6a2fef028f935ace3

Observation 4afe23c0-593d-431b-9087-6b094fb3db7d · outbound

This paper cites Beavertails: Towards improved safety alignment of LLM via a human-preference dataset.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Beavertails: Towards improved safety alignment of LLM via a human-preference dataset

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.734394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.040891Z digest=sha256:c2f006baac2870d403c0d6e1f28aa762d96c1757d13d93528027bf410f2bb8ab

Observation 3de26fcb-5bcd-45b2-afd1-e7c91484d7ce · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.043782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.043782Z digest=sha256:4e6c2b24aaf021daf6c4b83215f6ee8a6db80ea0fb387f9534649b20faf55e11

Observation b2953044-edaf-4691-82c5-bebcce94a697 · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.046987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.046987Z digest=sha256:079fd6d3f3039fb0bdc26e914a1cf7e1cc0eb675ae20a1026e2cab5c96ee7455

Observation dc250669-7669-49bd-8277-958dea84a5d3 · outbound

This paper cites Controllable Text Generation for Large Language Models: A Survey.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Controllable Text Generation for Large Language Models: A Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.050480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.050480Z digest=sha256:8a4db129e073f346f2dbea55fd9772c0aed2107e8e1c9b4e613b3e2dbb9ba87a

Observation 3c54c4e0-4462-4e33-8941-af2e662b11c2 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.053735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.053735Z digest=sha256:8c1a7d4fcdd4b6a540cc5d038f24a606ce4bb9f97d98ca942ab221cf4d75ff60

Observation 7c9328f6-2725-4373-b282-8c910e68d096 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.056797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.056797Z digest=sha256:5df88f486cbae731ad31d50b36370630c0fc2d34b449f70dc7e580fa1fb3947e

Observation 319a4ab5-96ea-43be-a1f6-df6a5040b93f · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.060395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.060395Z digest=sha256:320736b2c896f8178b0aca0ee9b894a31aacfafd55c641a2ee559346e184ecc5

Observation 8a1a9168-3123-4793-9419-953e9e725e33 · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.063868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.063868Z digest=sha256:039cfdf687f89e5db56bca58080e40b7c5c235795a60910dc0e9a59987fc8bf4

Observation 07339d1c-e2e0-412f-af1f-cc06439acf89 · outbound

This paper cites Meta llama guard 2.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Meta llama guard 2

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.725163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.067238Z digest=sha256:e3e1af36ff13b85c337e13a25750bff167f7f140db7727ff3acf71d4ef7a19f4

Observation 5ff434a2-881c-4014-822f-0c9e403eee70 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.070762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.070762Z digest=sha256:6cb1786a9b6a997e0d805f975212ea98a40aaf03629b691120c00583109be0ae

Observation cabeddd7-716a-4913-be82-5143ffd63a3c · outbound

This paper cites Controlled decoding from language models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Controlled decoding from language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.716141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.073861Z digest=sha256:641ec96ccf74777779f7bede12082498b02d67cf5090d015f371d29fdd33c378

Observation 45c92630-c3b6-4cf6-9000-1558fc446376 · outbound

This paper cites Training language models to follow instructions with human feedback.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Training language models to follow instructions with human feedback

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.707231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.077142Z digest=sha256:3933a8632d05a189c89faf3713f1e1fb85c90281f956f8a963ac8d42b40e72a7

Observation c2952507-47d7-4919-8b83-911c110e49ec · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.081017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.081017Z digest=sha256:449b95971289cfe09b815c5c06e76954fe91a75fcd20853f12da33f663c83599

Observation 3bfc7c9d-2f7f-41eb-b1cd-3363dd6a452d · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.084092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.084092Z digest=sha256:cf883512684683e61a87fc7d1c40675288139e6783d1b0c8a95ebaa06f060d8c

Observation f0ddc9fd-24d8-469a-b31b-710ddb4b2787 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Direct preference optimization: Your language model is secretly a reward model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.087213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.087213Z digest=sha256:ba83d1f896cdb922674e1a7b91268bd1cd064bae28e3e2b6b7cd471b782b57cc

Observation ceac96ca-22cb-4c6f-af82-1ec1d217ac9c · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.091162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.091162Z digest=sha256:d6b41a33ff01d16d663a9f97a76a5ce5d858d2ac265be8e4cc7f427090dc0e14

Observation d36f4813-4d84-4cbe-9bed-3d4848a17769 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.094349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.094349Z digest=sha256:7db5ce1b18f708fce8f46d01a6f3e33107da7dec898383b485d476a7ccb0f221

Observation d2fcfa10-a98d-45f9-9406-c10bcee36b09 · outbound

This paper cites LM-Nav : Robotic navigation with large pre-trained models of language, vision, and action.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing LM-Nav : Robotic navigation with large pre-trained models of language, vision, and action

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.691469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.097713Z digest=sha256:2a7b213838b802e7613a08429f51505f3e93c55cace0379e92350237b66a4fb8

Observation 1166d32d-20bd-4c3b-b1e8-4d4ce1f46fc9 · outbound

This paper cites Hashimoto.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Hashimoto

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.100760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.100760Z digest=sha256:e3f907cf4e7f459a9d186a7389e60452663240e9f8bd140bbfd2e6db5eb6aab9

Observation 01242aff-dd32-464f-872b-43b3309c24e5 · outbound

This paper cites Large language models in medicine.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Large language models in medicine

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.104154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.104154Z digest=sha256:9ae1213d978450df795372cb88e5cd918b840ada588336cf37263b25811ad668

Observation df906c89-96cd-4fb8-a106-c46f4b520c8f · outbound

This paper cites TRL : Transformer reinforcement learning.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing TRL : Transformer reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.671249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.107392Z digest=sha256:cf286185e0f0531453e6ecaa981b84552c4b607135a8744248406e37701ea042

Observation 306839f3-de22-4a95-8c25-eb9eaf73ecb4 · outbound

This paper cites Stepwise Alignment for Constrained Language Model Policy Optimization.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Stepwise Alignment for Constrained Language Model Policy Optimization

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T13:14:34.276505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.110335Z digest=sha256:540029b639cab887965b72984e03d3997c27581cf87d548bd865f64d31e9b0d9

Observation b219fbce-6ee7-4844-887d-2171c5551971 · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.113399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.113399Z digest=sha256:96f78fb8990bb564f1ffa83dbacb9f0d1551cdfcaf07d5b3adbbb1500b848bcb

Observation aebf8f1a-3c1c-4fe8-a5c7-c90c725ec6d2 · outbound

This paper cites Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.116590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.116590Z digest=sha256:ab691083d954c5d11f8135e153bbc79580a39f7a1bdc091505f1c7d3d678d740

Observation 7bddb868-48fb-4ea8-b6c4-115df35357a5 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.120042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.120042Z digest=sha256:581a354efd64d76c2284ec2bd1b3073cc56af014964d67fd7be49cc756e007df

Observation 8161e2b2-6649-4b2c-8fa1-387350c8f91d · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Defending chatgpt against jailbreak attack via self-reminders

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.656375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.123076Z digest=sha256:f070ac6ad08a0fc56a07c7755992cccc38df11a2a654882566d09ff96ace1ead

Observation 2ad5719c-32d5-46eb-b9c6-9bcddfe21f03 · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.126221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.126221Z digest=sha256:ebdc98f0f7be1454045337e63bbf027310eadca6dc1c35d8a02ba98b55f141be

Observation 36981ff1-9520-433a-ad4b-23e1a8cb120f · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.129562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.129562Z digest=sha256:143146b25658c68fd7f94afa58afe21e0ecea2a6e2978bb09cc5f36ca2829f3a

Observation 81b9a09f-61db-463b-8b6d-d0322382fd37 · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.132795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.132795Z digest=sha256:55b4aba4baea0bd992802bda1eb97c08edf07c9af606f143251952da07272720

Observation 964664c9-7fef-45f7-92a2-028c2ff6a4a3 · outbound

This paper cites On the vulnerability of safety alignment in open-access llms.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing On the vulnerability of safety alignment in open-access llms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.647067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.136025Z digest=sha256:5f74a965f24ae4dc9ecba551def773cb14026135aaeb41480d1de44d27a1a0db

Observation 7c1fbb45-dc78-4bd2-b6b9-a3ed47dc4c18 · outbound

This paper cites Wordcraft: story writing with large language models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Wordcraft: story writing with large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.637485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.139018Z digest=sha256:8b1555663a249f8f2c2f9c2bbaef9e6975d0323ddb04abbcf50c2489eaa6e23c

Observation d7069d77-86be-4656-864c-0050591e34ce · outbound

This paper cites Prompting large language model for machine translation: A case study.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Prompting large language model for machine translation: A case study

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:14:34.627614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:14:34.141806Z digest=sha256:da1989592d06670d7de8bcca11754ed7584fe86906180fb40f1dec8270029ff2

Observation 5f3243ae-972c-4be3-991e-2052a8ea7f29 · outbound

This paper cites Panacea: Pareto Alignment via Preference Adaptation for LLMs.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Panacea: Pareto Alignment via Preference Adaptation for LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.144946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.144946Z digest=sha256:30b6297b6ab9b781d55e826eb4a659c1c966f9945bd642babf5ad46db56ea76b

Observation a96d3a07-5cec-41bb-a301-0f48fbcc06ff · outbound

This paper cites Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.148019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.148019Z digest=sha256:9a6cc012c606800d38aa438e479227a85678fff1fd0e36b19e22981545c8200d

Observation d7d0fdc2-7580-4405-948e-bdcdcc2e787e · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Fine-Tuning Language Models from Human Preferences

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.151110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.151110Z digest=sha256:f050085081eaeff21777a125ad6a0737b4f2bb7d646a0d5e5eb17387e2febe55

Observation 24ee1b5c-135b-4add-a6e1-7b061bef9de6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.154654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.154654Z digest=sha256:6d840ea7c38af8ad6151b6b550982555e5e0be0dc3ccbfe91a89aa302d83f4a7

Observation fdba3956-662d-498f-b783-e492e8e943b5 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Improving Alignment and Robustness with Circuit Breakers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.157748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.157748Z digest=sha256:4f582a728c9f991a0bff92a7507e538f3cab0ffc5a51dc4fb45278ad83de0141

Pith citing papers

No inbound Pith citation observations are available.