Pith. sign in

Paper Citation Record · LEDGER

Safety Alignment Should Be Made More Than Just a Few Tokens Deep

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2406.05946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05946 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:47:24.461017Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 18e07020-4633-4cca-ac77-ddc13a509810 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.408949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:19f7bfac49fb62074df5aa50c4e7fc369a746df044def30755cc4893a214cd4b

Observation 0a3315e6-5034-416a-a3a2-574c9ae1d9e1 · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.321187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:be6379aa695e0357bb2f31473fb5b61f3d7750799a389a859c0232e3aa34b011

Observation 69257a21-2ff3-458d-9a67-ac2adfc991f8 · inbound

Towards provable probabilistic safety for scalable embodied AI systems cites this paper.

Towards provable probabilistic safety for scalable embodied AI systems Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:02:15.176084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:00:27.799347Z digest=sha256:5c4373eb098230373c027ca3075eccdb7c48d306685dab3e30e66c8aac610ca0

Observation 31a3c0fc-9898-45f0-8b3b-81da89cdccf5 · inbound

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments cites this paper.

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:02:54.865061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T01:02:07.088724Z digest=sha256:19ce09a6db5b810bdd5b79875392e83297540ecc0862d6a3f430acb506998d74

Observation b18694bb-f4ed-470d-95df-23f3ff646588 · inbound

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment cites this paper.

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:47:24.461017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:47:24.461017Z digest=sha256:605383921d961d3e9d0a73d44248e51b0a7068485951206f39a9095da5c76011

Observation 9a85dacc-f99a-4773-9cb3-851478a24a55 · inbound

Reasoning Up the Instruction Ladder for Controllable Language Models cites this paper.

Reasoning Up the Instruction Ladder for Controllable Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:07:54.996168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:07:54.996168Z digest=sha256:e37c1b6155cb0930474f6794e6002fc741e54a45bba2feb070deab057eebfac3

Observation cfbb6043-d7a0-4edb-9bf6-cf0e2bbfa64c · inbound

SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models cites this paper.

SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:28:40.713873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:26:48.405593Z digest=sha256:fe453589c43a37838e3f156948cc2395ee1f33def98ee44bd16a9fdef97a4282

Observation 796e4efe-4123-4436-9ed9-85fff8302d09 · inbound

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control cites this paper.

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:29.072310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T11:17:03.104902Z digest=sha256:58849e0c7b799952c79cbb27cd86cf2d638dc795444b55b16c475d8774734e96

Observation c534c6af-9e58-40e0-bbba-33d68b94a85d · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.169526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:8ce494cbf010c33b37345250055b91adce0e4a88f5b4755999985363f59b6b33

Observation aefd7fba-2534-46b6-8ec3-068284866237 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.662413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:cc0971ebdb6323b9546e3b78533ff6e19414062c989c7d8e73c038635cc0e85c

Observation 3731c79d-7bc6-4d28-9d1f-c74def9e77ad · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:7a9313e458df52d0da20320203c2fb5555cf28122a3b055e1b9fbe2b259b6e7c

Observation a2dfbc05-f563-47c7-aa30-380075ba452c · inbound

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection cites this paper.

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:30:52.978026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:29:27.373741Z digest=sha256:93fc21c9b01871715364ba136830663300f8078ae14ef6d8f3aff202833a3045

Observation c548783a-0369-45e7-9738-d533958b6ae5 · inbound

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types cites this paper.

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:31:00.406851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:08:25.471462Z digest=sha256:55f04d102d249cf4e297e8040ad9d042531ef6c0e9dfda508599a9f994daa91b

Observation 5f398ca0-4b21-4259-9724-775fd701064c · inbound

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs cites this paper.

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:21:00.097938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T16:41:52.440793Z digest=sha256:c18327181404253d12314589a6e96ff7d79bd53bbac8e1419672f35c7f228f81

Observation ccc0fa3b-dfe9-4e8b-b4b9-e7df34205fd6 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:00.975345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:ae61e160ac88b54ccf2a5610862ab5990b954a1aa5270cd10dcf3247f612f8ce

Observation a2869571-ad0f-4ee1-8546-bbc2319adc3a · inbound

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs cites this paper.

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:18.735027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T19:22:00.217729Z digest=sha256:a297a0de4aba05a7dc85d6b41ee74592ce30fdcfc29e216e8991ebdca6b835e0

Observation af70cf46-cef4-42b9-b766-97bcd16061eb · inbound

Internalizing Safety Understanding in Large Reasoning Models via Verification cites this paper.

Internalizing Safety Understanding in Large Reasoning Models via Verification Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:51:14.234952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:50:59.283409Z digest=sha256:fdddb315b22b40e9569a4058add158e6b5883dab2873743eba997a3165183c0d

Observation f8c0cbb3-a741-458e-9e23-509d52f38403 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:19:26.730208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:3bbeff180e490c37ee6c45b5c2e5727ff188ff5dad02349b5937062c8af1cfc1

Observation b8910e1c-87c9-4d1a-ab7a-1ba54351b815 · inbound

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation cites this paper.

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:39.886052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T16:34:47.856606Z digest=sha256:d9b72d854b076cf56c919694bb9748a071900228bc14eada12baf6378c4738fe

Observation 50336ac5-2974-406f-ac2d-16be30884918 · inbound

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications cites this paper.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.534733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:073343d8b549f0d1727b7aa443b5b6d038b83dd2ecad5c60331a539f367b7fd5

Observation 32b70579-c676-4e14-9cb8-9813176209ca · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:34:02.553168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:594b60f72186b2989fac2e53eb98d775400e24201b5986fd845e8992d1be3dcd

Observation 0c86dbcc-5cd4-4e1c-85ed-99cd3007d537 · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:01:19.405588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:4f69a87eece5dd9bdf540815c83d59a27b98adb59e1cac070c2e40c7563aa5d4

Observation 45b8de59-d327-4b5c-a66e-34520b9d6e45 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T06:19:41.984616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:537c3d4c87f5e1d358af6cdd466fedfa38814d1af698aa2e49fefd05a2166b9e

Observation e74004bf-4bd9-4474-990e-04ed45a96364 · inbound

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy cites this paper.

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:53:59.549110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T21:47:17.894881Z digest=sha256:7584f346c08c0c29cd131736666846f6892b4efe7e0933c27414ca86aea6db6a

Observation eeb1c9dd-8132-4ef8-9bf8-8e403c29f35d · inbound

MESA: Improving MoE Safety Alignment via Decentralized Expertise cites this paper.

MESA: Improving MoE Safety Alignment via Decentralized Expertise Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:35.291030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T18:52:00.377915Z digest=sha256:00d7388b4727beeccbedfad44e85c9b61c0f495122c3409a20acc790a4ac52e7

Observation 54d52469-bcff-45c6-be6b-ec2b969b7c84 · inbound

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models cites this paper.

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.704818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T17:23:52.388431Z digest=sha256:e3148f905db66c564f89b45e13fa9272c8f246bdaccbe1b9978d27810ad2b577

Observation 40e2d7f2-51b7-4498-9ae6-418d195d0f72 · inbound

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models cites this paper.

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:56:24.322243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T13:43:51.171443Z digest=sha256:d4f1f21f590b218e5934d237d796b4217dff048472d13f6a75d51c1dc4a89260

Observation d974c386-bb2b-494d-bd0a-1181b7910d2b · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.042879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:d2b8d207542ea97b72fa16aa7d9912ae3b74249a890591e504943c98f0c3f0d3

Observation db63a811-404a-41be-84c2-f94e3316018c · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.136806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:f85cccea70c65a1b683b01c86d64192fe3a5e136f6fd4ea1cede938772a5d653

Observation 3509ee17-abb0-4704-ad67-20241ed6968e · inbound

Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications cites this paper.

Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.668111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T20:47:36.337189Z digest=sha256:1d1c7e40ad34c2548b11bf668bd72e3d9d293a9a564cf46532c9d47386f237d8

Observation d4d13048-04c3-4378-bfaa-27f6348b12c1 · inbound

Forget, Anticipate and Adapt: Test Time Training for Long Videos cites this paper.

Forget, Anticipate and Adapt: Test Time Training for Long Videos Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:52.982067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:34:05.109347Z digest=sha256:32707a361aef200c4748c3702f02961881939210e6555dfd24c443c85f2d5d17

Observation 822981c2-9a40-4540-a0b8-1fb41754b31a · inbound

Forget, Anticipate and Adapt: Test Time Training for Long Videos cites this paper.

Forget, Anticipate and Adapt: Test Time Training for Long Videos Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:54:38.971589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T10:22:49.075651Z digest=sha256:d7b9f9319aa517caf5b3f190231a4eea203ce17e07404a37df4c4a648743d5d7

Observation 72b76032-4b80-44a6-9452-18ff921c8d81 · inbound

Forget, Anticipate and Adapt: Test Time Training for Long Videos cites this paper.

Forget, Anticipate and Adapt: Test Time Training for Long Videos Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T17:16:44.626166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:16:44.626166Z digest=sha256:875b75aa865b0f934244bb6ef596f56ca6b57d3496396243d730cb86dfec6b23

Observation 08176c57-f6af-417e-bb42-155045e9ae68 · inbound

Defending Against Harmful Supervision Hidden in Benign Samples cites this paper.

Defending Against Harmful Supervision Hidden in Benign Samples Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:24:45.328038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T05:29:04.802064Z digest=sha256:56cb510ac5804ffb09f1154feab444d75ef114a89cfd277bbd14808ff5ed85c4

Observation 52b6147b-d107-47f2-99c0-493177db0293 · inbound

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment cites this paper.

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.146959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-04T01:17:44.209685Z digest=sha256:2a9a66cc79749efd273d989648e3e49bfa117e54e4f0af8dddbb2e219d7ecc89

Observation 13b80043-0285-4561-8a97-cda8b8ec98cb · inbound

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 cites this paper.

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T18:20:40.287006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T18:20:40.287006Z digest=sha256:c6b33b99c6816043c70dd022f955525e6ab24c99dc97d162392b3fbfb4c1fe62

Observation 8b0e5569-2443-405d-8e81-3555d29cf70a · inbound

Pretraining Curricula Enable Selective Fine-tuning cites this paper.

Pretraining Curricula Enable Selective Fine-tuning Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T12:36:24.747752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:36:24.747752Z digest=sha256:d5703c01d5c3ff5010ad72d6d9f92ab24cf6827be6811c96e1572cf3f7cb3ba1

Observation 3df96879-e39b-4505-8a87-b1cbe822d585 · inbound

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? cites this paper.

An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon? Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T00:42:24.432562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:42:24.432562Z digest=sha256:215edc67d5ae2b5af12902c7d9825e4ed83b8c34928662572dd6283a68ddeef1

Observation 54e89284-a98d-4c11-8add-40007f5d5d99 · inbound

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak cites this paper.

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:30:56.251349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:30:56.251349Z digest=sha256:99ce1beba0da6233416a40f096204e20900649f525d0c48f7d52c7438f20a2a1

Observation 4c5b5df1-bb7b-4950-8fc5-89572e115e3b · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.569910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.569910Z digest=sha256:f87c773f204204a1283f6898e5a89e8872617ec45de8bbe07ad86c428948caef

Observation e94f6a3c-d03e-4e58-8f17-d7a4b741ab86 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.776258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.776258Z digest=sha256:adf36ae756a9cfbce1755d133c6e5da4ad8a8a30f156364bb6e771052815cc5b

Observation 18bb393f-b765-4321-a81a-61f84477da3b · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:54.536465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:54.536465Z digest=sha256:a512f4a975c5b88cf5668762f722bde4d64a148fa46227ccdbabfe3c9427115c