Pith. sign in

Paper Citation Record · LEDGER

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 6 inbound Pith citation observations for arXiv:2507.04250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04250 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:57:44.976131Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T16:03:12.728352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:29:39.399032Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d3d4fb5-83c3-42de-a7a6-611ceb89fe6f · outbound

This paper cites Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.379859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.379859Z digest=sha256:3db2b31adfbc66f271479d6c744df989691d67bc705daa7edf86f75b36ebbb4a

Observation ad634ece-212a-4348-a4c2-84b347ce5ad0 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.420384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.420384Z digest=sha256:4893c8ec604d00e1498acc9f122705142d6e5be7c32df9d0bf5e61405a85e5d7

Observation 01d7e062-d6e8-4049-bd09-02442604577d · outbound

This paper cites an unresolved cited work.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:57:46.335341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.505443Z digest=sha256:4f9ac36e9fd23e83a6b0e85cb858b78ec5e4345a99ba403c42723acde90cb50c

Observation 6363d2f3-5563-4c94-9ffc-1ed0b9af19d9 · outbound

This paper cites SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:57:45.412418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.549520Z digest=sha256:9d673f3f9cc81fd76f6cd17f16455dc03b679e0004632d6fe0b38e2046c306cb

Observation 245fc8a6-5db8-4ccb-8ce6-e13192ca047c · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.629886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.629886Z digest=sha256:f85fe7957947fa175b0635686c62f34a78e3c7ecfe74cd20a8dd63a3f1d1897d

Observation 7c88bbe9-127b-40f4-9ca0-8bab13ac7b53 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.694438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.694438Z digest=sha256:20edd89b63133572444715050db793787c51b585ce8887741ee5d2b7a65f54bb

Observation 63cb46a6-5c23-4a45-ad36-46c5ebbd73f1 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.738959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.738959Z digest=sha256:1c476ea14dec7aab85affcfd1132a85e8f12bca8fafffd37306e10bdb266a4e0

Observation 3159a8a9-bde1-48bc-8dda-7e11eb0f4802 · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TrustLLM: Trustworthiness in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.796231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.796231Z digest=sha256:1db9ce8de474bb89e2b4fd5c8989acd3f569ee515398166aaecfabfbb9cf9791

Observation 7af4eff9-96b4-41c6-a3f4-5c45c0eb443d · outbound

This paper cites GPT-4o System Card.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.829835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.829835Z digest=sha256:2c40a859b43be8c2d1cd11774943e45a94c8ef47417aeb44138bb19f0b0685ce

Observation 34744acc-9d3e-467f-8461-8daf47ba30b2 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:46.158196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:57:43.856574Z digest=sha256:9e5454bd0f192437d1f64bd13b0f9702b67e42b25f9ff89ed924be750d988765

Observation c5e47030-9380-439f-a9cc-566b8510edc1 · outbound

This paper cites Crafting papers on machine learning.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Crafting papers on machine learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.914249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.914249Z digest=sha256:85553cfe234ab9f55cc078fd1589b6978ffce304f78d3c43f44092fbe4ffc95f

Observation 7cc02672-ba13-4cb4-8174-ab22e8cf3187 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.958573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.958573Z digest=sha256:8d9dc1b95d04ffa71e1b7b80547fdd6ea8b2d2a15dcdb03f4a554a3b1d920216

Observation d8f4e40e-329a-4c5d-aa15-0a3e987f0f32 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:43.979623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:43.979623Z digest=sha256:2dab32d71d32050ca56ecc97d139fc4d2ccb15eb26cf24dfef7405eb10bc887e

Observation 6f78b735-0c61-4eb6-91af-9259c8a13a8a · outbound

This paper cites Decoupled Weight Decay Regularization.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Decoupled Weight Decay Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.013076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.013076Z digest=sha256:d99ad95aa863ed1f926e8afe1081dc380b40a369883d912a76a7d23b47030c26

Observation d719852f-f4d1-4777-b158-d313ce2019ae · outbound

This paper cites Pointer sentinel mixture models, 2016.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Pointer sentinel mixture models, 2016

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.069618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.069618Z digest=sha256:a606c750543f085c8a3f76f8a3303e157b47bdbd528ed914e0d02b473e4c8d2c

Observation d2b9753c-600a-4924-b2f5-462fc90e12d9 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.977389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.107947Z digest=sha256:0c7500ffb391504e6b93e9b6f6a93b9ecacf5b8ce1b2a412d46509b550c6c875

Observation 3a62f8b4-c2ec-41f9-992b-39d757f069fd · outbound

This paper cites Mitigating Exaggerated Safety in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Mitigating Exaggerated Safety in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.141968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.141968Z digest=sha256:2830e508deb5990c97a89a7db083680bb4b5764b8ae845540eddf4f4d74b41d6

Observation ced8ca4f-2b94-4e70-9c0a-ba241d4ad76d · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.190147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.190147Z digest=sha256:3dd69cbd3d0840bd025f6d5ff9de6c5d3ade5fcd9408eaedf3cf00b46be4d8fb

Observation c75906bb-a888-4fc8-8a57-a2d0f659bbe1 · outbound

This paper cites an unresolved cited work.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.251439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.251439Z digest=sha256:8e0d490c49cb9bb0c383d47ec4abfe95ec29a418af393841d741943d46889e5d

Observation b9e105cc-1196-42d2-a5d7-6a6921395e5f · outbound

This paper cites Navigating the OverKill in Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Navigating the OverKill in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.284799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.284799Z digest=sha256:fbae8cf58f874d07dc2ed519399c6252573ec732f758b7bb9d020dea94cdcb5d

Observation 0d864902-e309-40c2-9228-8090fcdddf14 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.329662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.329662Z digest=sha256:ab31db6b3faa8bf28b6ca120e4794ff13a3384d134ff9dc9d9407c5294344a4e

Observation 9f41f133-8e1b-42a4-96e3-fe5307c65c16 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.388456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.388456Z digest=sha256:7e96850f3b58d248ea9e79c0d08d750d0c526092d87f65d9553b71100837402d

Observation 113f4a0c-e4e9-4fc3-8c43-9b0ccdfa1a4f · outbound

This paper cites and Hinton, G.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning and Hinton, G

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.827219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.415871Z digest=sha256:ee9877ad4fd9ae33616bdeeef4235c66888055315b4813281d518b66e762b55d

Observation 003f4759-e069-41a3-b2a5-b6a1e2628d3f · outbound

This paper cites Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.461260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.461260Z digest=sha256:64b2b4b4de604e23a09a6c3c5770d1046e179f5142278084a6d793d88af8393a

Observation a418dee3-7334-447b-b896-45a9f0069e5e · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning ReFT: Representation Finetuning for Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.494978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.494978Z digest=sha256:ffbe0fab1591e0a7c1f97e40bc280717774ae72ac6d9ee0a73fe11fd6cf3699d

Observation 6063ade9-b576-467b-acb4-3cb2196d7f7e · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.540195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.540195Z digest=sha256:630b5d07eab520fcff720a29220c5b04a1a5282e270acbb3707f8f4d42afe32d

Observation 59991026-be87-47d5-9ec8-ebff5c8296be · outbound

This paper cites LoFiT: Localized Fine-tuning on LLM Representations.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning LoFiT: Localized Fine-tuning on LLM Representations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.575405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.575405Z digest=sha256:88b2a954fa8d1cc0a7e6debc61625dfe3ddb7a30c069328cf8cc6a4ea8cbbffb

Observation 2da464bd-d208-4966-a920-d39ebb5d9d92 · outbound

This paper cites Scope: Scalable and adaptive evaluation of misguided safety refusal in llms.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Scope: Scalable and adaptive evaluation of misguided safety refusal in llms

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:57:45.642516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:57:44.646800Z digest=sha256:b88d7778a992befb253d44d968b94cb78f75dbc36addbb858eade92d266a0df1

Observation 53efe874-0171-43da-9d23-0d106ff907c2 · outbound

This paper cites Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.683011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.683011Z digest=sha256:792c4edf72941493d1c5f20125b20ebd11c20fb2bb62f2655d05130baa24e720

Observation 36cde0a9-52cb-47df-8885-9a0c480d83cd · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning On Prompt-Driven Safeguarding for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.717885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.717885Z digest=sha256:396c320b18c6a8b5729729a53b3ce273301fecde2d230c0bc5dcafbd4c35a983

Observation 656ec8c5-a639-4a8d-8595-528d21bac811 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.785662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.785662Z digest=sha256:fb747063108888e5a3f5f0cdda08147d54f540345afa3f4c799dd2be40cd6292

Observation 5fe881a2-87c0-421d-929c-8c8ad971607a · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Representation Engineering: A Top-Down Approach to AI Transparency

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.836509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.836509Z digest=sha256:30abb604f6e49c06fcf2f2b2502f3bc8b3832e7000a9149752e1a6c1eb3a12f1

Observation af12dd0b-7f73-46b2-b5eb-5f4f0ab62693 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.883812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.883812Z digest=sha256:7682179408068c0b0ac1a6e709b20b97ba0e5f816f789232197d278d2ae8b225

Observation e505e4b7-5dbd-4cf7-9021-f49fb8640126 · outbound

This paper cites write newline.

Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning write newline

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:44.976131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:44.976131Z digest=sha256:261095a7a00e5c6c1964ecd478d7d509267af9bcb3ec1645002dc6272332f454

Pith citing papers

Observation 5dffa089-944e-407c-9c27-475a836514af · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.523835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:6d85d6b73f8922d1d3edbda449854e4065d862a478d1911b7758c1e53b3e5e39

Observation 27eaeb93-936b-4df8-99e9-37edfaa29333 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.385931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:874734abf30c3a7d80df85ba7e09f72f304870f6c0105434b4dd8802fa2b379a

Observation 155ce5e3-1110-4b05-9f61-83d5fcff2aee · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.856739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:65bf5fd18e7e532adac8f3a880b34f9af171138903fbf5e2634e10a2ef445dcd

Observation a1ba78f2-a143-49b2-b127-53827db5d500 · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:04:52.603163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:0e60b823ee9de8b247c0252d893be542ad94a605c127d72741e51b9b935951e8

Observation 04bef1af-3585-4424-8c72-813677dafedb · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.873093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:a145de3c0914985cf8dfe540cf4afde131c495de02fbae0086bf8b45d8b88fed

Observation 7a11ef31-8b4d-48ac-a05c-09782b7accf7 · inbound

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? cites this paper.

AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:39.400611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T13:22:12.541923Z digest=sha256:6e8686020f5244a3038fe86af6494b43da6cfa7babd9b69eb9d87bbdd6994ec9