Pith. sign in

Paper Citation Record · LEDGER

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

As of 3 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 1 inbound Pith citation observation for arXiv:2604.12371.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.12371 v2

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:37:19.310821Z

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T02:44:13.401198Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4ea2ad9-f244-488b-a5b1-93850ff9909e · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.685664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:3d1b99c5bb1d0e71359eba3de6b8458e180bad61a786be84471f60ec18e8616d

Observation 24cfb540-00e6-46f3-afe0-87e394e51efb · outbound

This paper cites Salad-bench: A hierarchical and comprehensive safety benchmark for large language mod- els.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Salad-bench: A hierarchical and comprehensive safety benchmark for large language mod- els

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:29:46.018962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:fda658b9d5008b019eae412903c60727e6691f90b7956e2531dcb250cd914aed

Observation 10b826a8-6271-408b-9c33-79d68de62f2f · outbound

This paper cites VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.681020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:337bccf74bbfd6abbbff60433ccd90ec16c58613bc781b7eb44f6ee258ade928

Observation 421ca7c1-2f1d-42c6-8f68-cfb678f97af7 · outbound

This paper cites Typographic attacks in a multi-image setting.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Typographic attacks in a multi-image setting

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:29:46.025966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:3101d71c97a599998aed392b776233f1916626a9d1eddffee713d4a29dd8516d

Observation 0978cd53-98d4-4bb0-ba84-f61af5960e47 · outbound

This paper cites SCAM: A real-world typographic robustness evaluation for multimodal foundation models.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models SCAM: A real-world typographic robustness evaluation for multimodal foundation models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.672907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:cb4f111531ceecb20ccd1ca2728742af4f004c5c0ebb0c3c2747363442a764ff

Observation 985309d2-d4a7-4328-9fdb-93f9c59ad627 · outbound

This paper cites Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning.

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models Advclip: Downstream-agnostic adversarial examples in multimodal contrastive learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:29:46.022527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:37:19.310821Z digest=sha256:db113d752ef9f48068b25c66e14206d9f8f28ec04113a303fabc2fbfeee7cca3

Pith citing papers

Observation 25ec6ad5-7450-430d-8bc1-36f5a9a20c14 · inbound

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation cites this paper.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.401198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.401198Z digest=sha256:dc2f11230fb76909044bbb0430e169af7e7b64002f51a5b1e6c2264e62c03767