Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:57:44.976131Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 6 inbound Pith citation observations for arXiv:2507.04250.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:57:44.976131Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T16:03:12.728352Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:29:39.399032Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8d3d4fb5-83c3-42de-a7a6-611ceb89fe6f · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad634ece-212a-4348-a4c2-84b347ce5ad0 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Refusal in Language Models Is Mediated by a Single Direction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d7e062-d6e8-4049-bd09-02442604577d · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6363d2f3-5563-4c94-9ffc-1ed0b9af19d9 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SCANS: Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 245fc8a6-5db8-4ccb-8ce6-e13192ca047c · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c88bbe9-127b-40f4-9ca0-8bab13ac7b53 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Measuring Massive Multitask Language Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63cb46a6-5c23-4a45-ad36-46c5ebbd73f1 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3159a8a9-bde1-48bc-8dda-7e11eb0f4802 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TrustLLM: Trustworthiness in Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7af4eff9-96b4-41c6-a3f4-5c45c0eb443d · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning GPT-4o System Card
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34744acc-9d3e-467f-8461-8daf47ba30b2 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5e47030-9380-439f-a9cc-566b8510edc1 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Crafting papers on machine learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cc02672-ba13-4cb4-8174-ab22e8cf3187 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Safety Layers in Aligned Large Language Models: The Key to LLM Security
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8f4e40e-329a-4c5d-aa15-0a3e987f0f32 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f78b735-0c61-4eb6-91af-9259c8a13a8a · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Decoupled Weight Decay Regularization
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d719852f-f4d1-4777-b158-d313ce2019ae · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Pointer sentinel mixture models, 2016
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b9753c-600a-4924-b2f5-462fc90e12d9 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a62f8b4-c2ec-41f9-992b-39d757f069fd · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Mitigating Exaggerated Safety in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced8ca4f-2b94-4e70-9c0a-ba241d4ad76d · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c75906bb-a888-4fc8-8a57-a2d0f659bbe1 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e105cc-1196-42d2-a5d7-6a6921395e5f · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Navigating the OverKill in Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d864902-e309-40c2-9228-8090fcdddf14 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Gemma: Open Models Based on Gemini Research and Technology
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f41f133-8e1b-42a4-96e3-fe5307c65c16 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113f4a0c-e4e9-4fc3-8c43-9b0ccdfa1a4f · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning and Hinton, G
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 003f4759-e069-41a3-b2a5-b6a1e2628d3f · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a418dee3-7334-447b-b896-45a9f0069e5e · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning ReFT: Representation Finetuning for Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6063ade9-b576-467b-acb4-3cb2196d7f7e · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59991026-be87-47d5-9ec8-ebff5c8296be · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning LoFiT: Localized Fine-tuning on LLM Representations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da464bd-d208-4966-a920-d39ebb5d9d92 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Scope: Scalable and adaptive evaluation of misguided safety refusal in llms
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53efe874-0171-43da-9d23-0d106ff907c2 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36cde0a9-52cb-47df-8885-9a0c480d83cd · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning On Prompt-Driven Safeguarding for Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 656ec8c5-a639-4a8d-8595-528d21bac811 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe881a2-87c0-421d-929c-8c8ad971607a · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Representation Engineering: A Top-Down Approach to AI Transparency
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af12dd0b-7f73-46b2-b5eb-5f4f0ab62693 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e505e4b7-5dbd-4cf7-9021-f49fb8640126 · outbound
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning write newline
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dffa089-944e-407c-9c27-475a836514af · inbound
Characterizing Model-Native Skills Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27eaeb93-936b-4df8-99e9-37edfaa29333 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 155ce5e3-1110-4b05-9f61-83d5fcff2aee · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1ba78f2-a143-49b2-b127-53827db5d500 · inbound
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04bef1af-3585-4424-8c72-813677dafedb · inbound
Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a11ef31-8b4d-48ac-a05c-09782b7accf7 · inbound
AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.