Pith. sign in

Paper Citation Record · LEDGER

Harnessing Textual Refusal Directions for Multimodal Safety

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2606.31876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31876 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:15:25.278988Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact7
  • verified fuzzy34
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ff4e6ca-df07-47cf-9da8-36ef48b65446 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Harnessing Textual Refusal Directions for Multimodal Safety Refusal in language models is mediated by a single direction

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.811599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:b78b6d842c64a8b2b7f83362190a78f7af3b390c83e2132f960e31ecf98eeb10

Observation f1980b79-27b4-4f58-80c9-d19b17722666 · outbound

This paper cites Qwen3-VL Technical Report.

Harnessing Textual Refusal Directions for Multimodal Safety Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.109242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:92f57fcffe31871c4468a867fda4de069f58457dac15af2863e9274b5ab44a05

Observation 7b3e5f92-4272-4550-bae8-de22dc7d646c · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs.

Harnessing Textual Refusal Directions for Multimodal Safety Emergent misalignment: Narrow finetuning can produce broadly misaligned LLMs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.770444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:92a1234e07a7127ab2cbc04b11e57577d7195faf5f48ef36092e29a2f73a647f

Observation 412ffa51-b829-4e43-a73d-ba9455716fc3 · outbound

This paper cites Discovering latent knowledge in language models without supervision, 2024.

Harnessing Textual Refusal Directions for Multimodal Safety Discovering latent knowledge in language models without supervision, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.796410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:8c9f0279aa4703ba944aaa14f759dee2556950d02eef4e1bf2e892b0285f72be

Observation 23c1a5a7-6ffc-4218-b1b3-4ab0328c9b52 · outbound

This paper cites Persona vectors: Monitoring and controlling character traits in language models, 2025.

Harnessing Textual Refusal Directions for Multimodal Safety Persona vectors: Monitoring and controlling character traits in language models, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.772243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:d6177154997768ca2a2d3c1d8aab7854b274a08f644dd811817abc737ffc3c4d

Observation 000a6658-f00c-40ba-9b38-0a5465b4c973 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural language feedback.

Harnessing Textual Refusal Directions for Multimodal Safety Dress: Instructing large vision-language models to align and interact with humans via natural language feedback

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.774095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:b2bcd3d55d1534048d05ca8a6badd75b0dcf6a214bc1b7f12cd676e67aa6ae3c

Observation c3270bca-9144-47e4-a0e2-fd637b55093a · outbound

This paper cites Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding.

Harnessing Textual Refusal Directions for Multimodal Safety Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.121467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:be1dec363fa97c931468e63b3ccf20cabf964ac84850d8ac9484cc0f6f1843cc

Observation f596339c-8671-45a8-a9c7-ba58c7bfd197 · outbound

This paper cites Gemma 3 Technical Report.

Harnessing Textual Refusal Directions for Multimodal Safety Gemma 3 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.116433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:85de4d6e78f4aced518df96194f01c440676cf7c8e4cbdb8d766bec105489a2e

Observation 64bc4634-0508-4092-aa6c-e9b4ec2e6273 · outbound

This paper cites Kwok, and Yu Zhang.

Harnessing Textual Refusal Directions for Multimodal Safety Kwok, and Yu Zhang

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.775830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:30146b5da6dc6333439051be18930390c41b863cd1d4e6f22e943cf67b30d60b

Observation be2bd658-f2f8-490c-8821-a72bdd96f1a6 · outbound

This paper cites Catastrophic jailbreak of open-source LLMs via exploiting generation.

Harnessing Textual Refusal Directions for Multimodal Safety Catastrophic jailbreak of open-source LLMs via exploiting generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.777701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:4c84dcfca4a7aa990e1b296e8643f161ffe5049f24f45a0644471824867cea97

Observation 1f1b017b-afe2-48ff-b604-30d7da9bd6fe · outbound

This paper cites Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre Dognin, Manish Nagireddy, and Amit Dhurandhar.

Harnessing Textual Refusal Directions for Multimodal Safety Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre Dognin, Manish Nagireddy, and Amit Dhurandhar

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.781336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:4e87acba8778e7894490032687701a76e831afc7b4f6d26e204723f44c8ff6b2

Observation 981a083f-c5ba-4828-b11a-0c40f2524b12 · outbound

This paper cites Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b.

Harnessing Textual Refusal Directions for Multimodal Safety Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.831268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:af4826f34d2825a0a8026706bf61db26a805b912b05530d6e57604f44da3c174

Observation 954bea8f-9ebe-4b8e-b785-91d46dfea4bb · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

Harnessing Textual Refusal Directions for Multimodal Safety Inference-time intervention: Eliciting truthful answers from a language model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.826539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:36d19744f3690611bfb4e2d81eccd15306cf5b371eb50c7ab3514fc0c5ce1802

Observation 336f2056-1a11-4191-bdd9-a2fbaf232eff · outbound

This paper cites Red teaming visual language models.

Harnessing Textual Refusal Directions for Multimodal Safety Red teaming visual language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.821002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:bc196462fcfdacba789cb2060925312cbf75c9c474e30b7f9fa95dc630fd9161

Observation b40283b3-17e3-4212-89fb-de8fbcbbd17c · outbound

This paper cites Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models.

Harnessing Textual Refusal Directions for Multimodal Safety Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.803679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:4071cf27bbf2ee3a6120bb25fc872c48a4bbdac9bbfd8b4a2858a42c7bfccb5c

Observation f2b30212-aab1-4791-9e68-5dc9b22030b1 · outbound

This paper cites Improved baselines with visual instruction tuning.

Harnessing Textual Refusal Directions for Multimodal Safety Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.798194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:6873fa289f7089b8b0a30f098b3e16a9ac5d7ca2f8678a6e746fa70952b4716a

Observation 021bf172-b99c-4bf9-b098-9a2622c3a269 · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

Harnessing Textual Refusal Directions for Multimodal Safety AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.817136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:4c3c5ecdf0864d2cd92e3a61d54d862f35ef8f3f886c0a2b2221cf27940d3b6a

Observation 8970adc2-db68-45d3-97c9-2afabf5ff638 · outbound

This paper cites Mm-safetybench: A benchmark for safety evaluation of multimodal large language models.

Harnessing Textual Refusal Directions for Multimodal Safety Mm-safetybench: A benchmark for safety evaluation of multimodal large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.819077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:334b88e82e860df45fd8d22f38a20d5d94e23b479736df4c29d1522c1a6e535e

Observation cc989b76-0944-4f5b-9967-92e7987e2a71 · outbound

This paper cites Video-safetybench: A benchmark for safety evaluation of video LVLMs.

Harnessing Textual Refusal Directions for Multimodal Safety Video-safetybench: A benchmark for safety evaluation of video LVLMs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.815239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:a0f0697ba60876981c303613c2c92ce5f6f3078660fa2dbeedb8ecc123442f42

Observation f924584e-c12d-4215-a2b6-0d5520c882a4 · outbound

This paper cites The Llama 3 Herd of Models.

Harnessing Textual Refusal Directions for Multimodal Safety The Llama 3 Herd of Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.123891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:c1db29a7cb8517e33902998e15a7df80931f2ce279c90ea445d6eee5fb16f6c5

Observation 73619660-986f-431b-a6ce-2cf630b9d8e2 · outbound

This paper cites Think in safety: Unveiling and mitigating safety alignment collapse in multimodal large reasoning model.

Harnessing Textual Refusal Directions for Multimodal Safety Think in safety: Unveiling and mitigating safety alignment collapse in multimodal large reasoning model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.829114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:4d2f98545efee7111991f3c4591d4c336b54b7937e54e35bbd0ccf55c3bc1d66

Observation 6e6e5da4-7d00-4b13-acf7-4216e6eae460 · outbound

This paper cites Harmbench: a standardized evaluation framework for automated red teaming and robust refusal.

Harnessing Textual Refusal Directions for Multimodal Safety Harmbench: a standardized evaluation framework for automated red teaming and robust refusal

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.785107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:3750811d74996a686bf5acbca0e201220c313d7a46eb8d5a8eb6edff519a3d8e

Observation 8c68beee-1317-4272-bb06-f168681dfc78 · outbound

This paper cites Red teaming language models with language models.

Harnessing Textual Refusal Directions for Multimodal Safety Red teaming language models with language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.789007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:b0448d15a43c11f2c3cce37ffe28e21793b46b703b249be78371c23ff94d8a07

Observation d092f225-6686-4c59-9490-af19c793fdcb · outbound

This paper cites Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models.

Harnessing Textual Refusal Directions for Multimodal Safety Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.790778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:e53b73c2b1df7f37de0d05dec3877cacf0662673a8bfa4271890fe3beeb60b1a

Observation 9b5297ce-4d0e-4124-a9dc-90fde6ded9ee · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Harnessing Textual Refusal Directions for Multimodal Safety Safety alignment should be made more than just a few tokens deep

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.794414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:e1eca60dc3b33eaea2f8194c574049853de86d1475e5c5adcda8ba55ae583bcc

Observation b1c612c5-b6dc-4850-bf6f-02053fbf2ae0 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, 2026.

Harnessing Textual Refusal Directions for Multimodal Safety Qwen3.5: Towards native multimodal agents, 2026

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.792490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:415b8b00408a537a0973f9c69788d67d3b145ac6e91e9d798f703f9f3e8cf67e

Observation e4bd932d-d251-431b-aaac-4758df47e3df · outbound

This paper cites COSMIC: Generalized refusal direction identification in LLM activations.

Harnessing Textual Refusal Directions for Multimodal Safety COSMIC: Generalized refusal direction identification in LLM activations

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.833160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:779a99d54ecac824a1ce28a0317232bbeaf7dc6c4bf135b74838289e9d85313f

Observation d3a14f4b-1492-4dc0-9ae8-c82cd155ab8e · outbound

This paper cites Activation scaling for steering and interpreting language models.

Harnessing Textual Refusal Directions for Multimodal Safety Activation scaling for steering and interpreting language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.779513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:4112d4656f842d1192395f80deeb740d8e3a107159943596cb494b61993c24c6

Observation e66d3783-96ec-4780-a366-ee800f4a9a91 · outbound

This paper cites Hashimoto.

Harnessing Textual Refusal Directions for Multimodal Safety Hashimoto

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.834955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:cb03cc25cd6d8e47add8075c8b112a783c3c9a27155749510e480fc7b1a2c879

Observation 991bcdbe-af80-49f5-b27b-8b896eeb976b · outbound

This paper cites Kimi-VL Technical Report.

Harnessing Textual Refusal Directions for Multimodal Safety Kimi-VL Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.118949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:d9b5b2609d5ab9030df146bcd31f483af685ec56937d61b3ed2a36f1abad8edb

Observation 05ccbb10-afab-406a-a260-f9e1d09ae0ca · outbound

This paper cites OpenAI GPT-5 System Card.

Harnessing Textual Refusal Directions for Multimodal Safety OpenAI GPT-5 System Card

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.111545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:d0fadb65ad91a50610363e2e3c74899b98bd0296dd697709f663b8d0dca8a591

Observation dc9ae0ab-7b1f-4900-8f6d-09e0a2d7270f · outbound

This paper cites Vazquez, Ulisse Mini, and Monte MacDiarmid.

Harnessing Textual Refusal Directions for Multimodal Safety Vazquez, Ulisse Mini, and Monte MacDiarmid

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.787052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:e8737848c52148f23013691eced30376d83c9ed4eb0ad1a00cb454f9a86d945f

Observation b094b0e4-616c-489b-9e94-af321c9ee770 · outbound

This paper cites Steering away from harm: An adaptive approach to defending vision language model against jailbreaks.

Harnessing Textual Refusal Directions for Multimodal Safety Steering away from harm: An adaptive approach to defending vision language model against jailbreaks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.813458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:d48b2969e4d2c519c5a611c893479594f261b4146e3f0ebe302cef1c1c067636

Observation fa8b1923-01b3-4351-a973-6ae9fd2b4586 · outbound

This paper cites Self-aware safety augmentation: Leveraging internal semantic understanding to enhance safety in vision-language models.

Harnessing Textual Refusal Directions for Multimodal Safety Self-aware safety augmentation: Leveraging internal semantic understanding to enhance safety in vision-language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.822856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:b4c10911bdf507b5396c9e6003a92c488b546be1356df6edb25ec04b95012920

Observation 5f77bc4c-9e62-4d09-8ce7-77f8457daf85 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Harnessing Textual Refusal Directions for Multimodal Safety InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.114280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:57dae0ca6b103156a5f2b21d7e6447bb9407f3c1d39cb4b726df6b04ae2b752e

Observation b820061e-9afa-4979-906d-12f3b862c103 · outbound

This paper cites Jailbroken: How does LLM safety training fail? In NeurIPS, 2023.

Harnessing Textual Refusal Directions for Multimodal Safety Jailbroken: How does LLM safety training fail? In NeurIPS, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.824636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:df9a1b82cf230cee75573187658e47703af2f10e40e511ac5b1d2a693fd4f3bb

Observation 4cb5f751-3070-4bf3-9ee8-44ed7a3abf3e · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Harnessing Textual Refusal Directions for Multimodal Safety Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.805657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:7430cbe08a916fe07928897d997013db9c0f1adc4697f20131e6d10a21fe0a3e

Observation c91697c1-9e33-4b42-b11c-a0916707c1bc · outbound

This paper cites MMMU-pro: A more robust multi-discipline multimodal understanding benchmark.

Harnessing Textual Refusal Directions for Multimodal Safety MMMU-pro: A more robust multi-discipline multimodal understanding benchmark

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.809342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:ab9581afc07955b55806a0c44a8a6adffb14606a199895d458fb9cf6072b05ef

Observation d249a07a-9bdf-47b3-9655-18b061911000 · outbound

This paper cites LLMs encode harmfulness and refusal separately.

Harnessing Textual Refusal Directions for Multimodal Safety LLMs encode harmfulness and refusal separately

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.783353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:c0074be596f5e25de576b1e5e398d13269dfa4463211e91a400276823bd0f28a

Observation 5d1c1fce-ac24-415f-9e0b-0b122e5095f7 · outbound

This paper cites Jailbreaking multimodal large language models via shuffle inconsistency.

Harnessing Textual Refusal Directions for Multimodal Safety Jailbreaking multimodal large language models via shuffle inconsistency

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.801853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:d4cff5b7a9dae685f76a0b5bc0e2cde0e0d0477ed287b87335ae15933b30ab4b

Observation e0542ba3-d181-4493-9752-794a42034510 · outbound

This paper cites The image shows a phrase of a product. List steps in detail to design this product.

Harnessing Textual Refusal Directions for Multimodal Safety The image shows a phrase of a product. List steps in detail to design this product

Reference 41

Resolution
malformed identifier
raw_fallback, observed 2026-07-06T22:22:58.800059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:b004ee7d9bfcd6746284c0cc79d9f74cc82c58f5ab99df9591328e5194bbed65

Observation 0eb35817-9356-40eb-adf2-a44dfdbb2f69 · outbound

This paper cites Candidate being Shot,.

Harnessing Textual Refusal Directions for Multimodal Safety Candidate being Shot,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T22:22:58.807465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:15:25.278988Z digest=sha256:be54a9446e6d81a50d3d73b68530154a576ad33d2383d8976b39b8b4ac91135e

Pith citing papers

No inbound Pith citation observations are available.