Pith. sign in

Paper Citation Record · LEDGER

Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2408.04594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04594 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:40:10.953813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T11:55:33.372242Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6e343e37-5f03-41bf-b747-fca572f1783c · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.925517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:95ba1105dc65ef34d2090f36285dad4694a64335ecb4b5e3ba88bb7f80b14882

Observation 2ef8ce1a-999d-44fb-bdec-d0ef11d773ad · inbound

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models cites this paper.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.953813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.953813Z digest=sha256:c9b81163d6ccda421ae552349cecd3177a90e6e10667bd848fe3503e29f4db6c

Observation 6ccdf34e-d9cf-4a89-841f-9bb7359e47e8 · inbound

Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence cites this paper.

Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:20:18.289420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:20:18.289420Z digest=sha256:92bb7c1d34b4054e65f0841bec7dc60aea4cd1dca87f79d2c03986255d058f52

Observation d45ef472-ecbd-4580-88e2-d8349dac5d33 · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.553252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.553252Z digest=sha256:1a11ed89506012648205ed169200aa7c15b6f57c9637f726eee23403f803c389

Observation 6ecfd414-2ba0-4de2-bf97-3d0dddc260f4 · inbound

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models cites this paper.

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:44:00.638413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:44:00.638413Z digest=sha256:1ec6d8ef26b429e29c63e969442e8e8261ba1b30a48a0cf248392e9b81281ec5

Observation 0a710233-2290-482c-8601-ea476edc3b1f · inbound

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison cites this paper.

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:55:33.374124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T11:54:18.587529Z digest=sha256:eed0217d889231706ad39dcc2da8b249819a7ec6fcc3401444fd875c0ef4a13b