Pith. sign in

Paper Citation Record · LEDGER

HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2502.14744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.14744 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:08:02.283463Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d041aa4a-cf62-4cd8-a29d-8289e8798809 · inbound

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions cites this paper.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.283463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.283463Z digest=sha256:4e9352790d8873fb23a4d0f05d54b6a4b44a9e71a0a60a4ba9cf51ffbd968a57

Observation 01b9a4ff-b677-440c-aa06-5bcd0b721490 · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:50.813253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:50.813253Z digest=sha256:f6ab8199a8cd50289c2d8460ff651b985646d53923c80c3cd16bda05752985d0

Observation 94b3aef8-1d9f-4329-9fb6-cde640ea49e5 · inbound

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models cites this paper.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:16.955465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:16.955465Z digest=sha256:aac857f5ab0c8640bd80dcbe45332977030b4fddb4ac38f61b1107637a11532f

Observation 229bef39-207a-4fa2-a909-ede030294c61 · inbound

Learning Efficient Robotic Garment Manipulation with Standardization cites this paper.

Learning Efficient Robotic Garment Manipulation with Standardization HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:00.311613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:00.311613Z digest=sha256:b49e8876c65b3396c05f7610a426b7301f9f60c95c1414b33013efda58c17b09

Observation e86577b5-cf63-4796-987f-4b919618607c · inbound

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair cites this paper.

The First Differentiable Transfer-Based Algorithm for Discrete MicroLED Repair HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:21:02.791999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:21:02.791999Z digest=sha256:2a850ae46959b45bc52b72fc077cf1a549a4b6c8c84cd13e9dc483ad733b49d9

Observation 7695229e-18a4-4341-b165-aed5a3105201 · inbound

Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception cites this paper.

Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:42.538039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:40:42.538039Z digest=sha256:92e488c23ce81ffd7ac2e8d42f23f9e324328e300d30e33b24123789f6738860

Observation ca1cfab4-bcff-43a4-87ba-f32f77a1ec62 · inbound

The Impact of Off-Policy Training Data on Probe Generalisation cites this paper.

The Impact of Off-Policy Training Data on Probe Generalisation HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.638926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:e66b55518920eb932d888f711299e6fb0677c609a63dcf1839763275ffaed804

Observation 2e4e9bf4-0386-432b-8f8c-e4ab437f8c22 · inbound

Adaptively Robust LLM Monitoring via Activation Watermarking cites this paper.

Adaptively Robust LLM Monitoring via Activation Watermarking HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:42.343655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:42.343655Z digest=sha256:f524db8815cc30bd7d490d91d0ddc3a53a820d9eb28f82f842d2823ed0104f65

Observation 50728afd-4e83-48e5-bd82-072f1a15b810 · inbound

SALLIE: Safeguarding Against Latent Language & Image Exploits cites this paper.

SALLIE: Safeguarding Against Latent Language & Image Exploits HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:51.149721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:04:46.426969Z digest=sha256:fd0a5d5391694a1b7c1cae2ed281e9e3533b633ccdc8418fa6e7c8e867f3a5f9

Observation 36de01fb-8a2a-4e07-9ff4-c79d16720069 · inbound

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses cites this paper.

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:46:31.207957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T02:16:49.785596Z digest=sha256:cad7b09eed6f2f5642a37770bb38ed1cb6b1b5ae81f524c294984ede161b3444

Observation a7b36e3c-27d4-4ff6-9aa8-3827602b0d25 · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:51:22.869020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:47:47.214726Z digest=sha256:ff2591402f517fe1d7c721620db54a07bce9aa680c319f3542eb785061e219e1

Observation 6b828383-b64b-4f1c-995f-7f81925d40d4 · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:40.487059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T17:06:22.481341Z digest=sha256:ea4d1b0757f78b16a2af317f1417a964ab5db7e867336df1239198438c3a1335

Observation a9da2504-4fff-4002-8acf-62ae90cc06e9 · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:37:14.918790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:750f67f2dacbb6d226d78f2bd3363909bfcdf42d7595ed7b8aaa882d0d687be2

Observation be50c1f4-ba70-4939-be24-6620ae2bb353 · inbound

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak cites this paper.

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T06:30:55.207937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:30:55.207937Z digest=sha256:d676f084092ec93c320ad0a67d87557768cf978937a23d4634f78aacd28cbde2

Observation 5b3c30c2-be6d-4e5c-80de-9472927ce71b · inbound

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure cites this paper.

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T08:23:30.045950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:23:30.045950Z digest=sha256:a4a396c3008295286dc964256955bd7a04a69ca39e8f2a5a187b5ccc3aeab78d

Observation 0e77d1f4-25a7-4ee5-865a-c23ebb2261b2 · inbound

MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration cites this paper.

MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T21:24:15.051214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:24:15.051214Z digest=sha256:aec5af552efd0fcd62389f4e262d7765d8eeaf9288f2cd85c9607988cb970125