Pith. sign in

Paper Citation Record · LEDGER

Contrastive Localized Language-Image Pre-Training

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.02746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02746 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:42:19.724749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:19:13.788988Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 10317e8e-dfcd-4da6-b919-f54e55f45b5d · inbound

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought cites this paper.

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Contrastive Localized Language-Image Pre-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:42:19.724749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:42:19.724749Z digest=sha256:c311267a8845518c30790a0b48111fac518039d76ff347e05cc073979c609451

Observation db769080-4e8a-4ec9-bb79-9fcfc021e317 · inbound

Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures cites this paper.

Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures Contrastive Localized Language-Image Pre-Training

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:37:31.659608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:37:31.659608Z digest=sha256:04c732c5d883ceaa27e73d9be7a28fc7414f3c25dd221943d42d4f02d2fe6154

Observation 97a0e847-3537-4f22-8915-6afb4ad9b595 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Contrastive Localized Language-Image Pre-Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.225148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.225148Z digest=sha256:66620bcc5ea014b673c34296c204d638771cd85e531e8237b6b32481cbe09f5e

Observation 41ed5110-7327-4e3d-bafa-1a1622ef2ee5 · inbound

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting cites this paper.

MedVista3D: Vision-Language Modeling for Reducing Diagnostic Errors in 3D CT Disease Detection, Understanding and Reporting Contrastive Localized Language-Image Pre-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:04.661105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:04.661105Z digest=sha256:08c110710ebc97c456a310ba35374b07e754c16e016b9d942e2a21d05cfd8a93

Observation 21d5ce15-807a-4a5b-b0db-72252ee48611 · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Contrastive Localized Language-Image Pre-Training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:42.682062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:42.682062Z digest=sha256:431c41b28239e63cc1570de847c8e4df3770d453c4452a68bffc15e163fc8a03

Observation 206d0ac7-b75d-4f82-bf73-4bb50d3769c9 · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval Contrastive Localized Language-Image Pre-Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:30:48.231935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:29:32.250418Z digest=sha256:b89986a8de4b98f48de59f850f9e5c5564d92d68d8254340c3e8757c645494ca

Observation b6db835f-6bf8-4700-a978-b909f73139ca · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval Contrastive Localized Language-Image Pre-Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:07:40.022288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:07:40.022288Z digest=sha256:74ebb0561f6c3f6eed6eccf2bd172030503d224e7d8b023527aa00c616d66e5e

Observation 14c2ff8d-487a-420e-9634-def8d146bdca · inbound

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning cites this paper.

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning Contrastive Localized Language-Image Pre-Training

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.562958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:16:57.329872Z digest=sha256:288ff22595cc37ec8708d607097958f01fca837691117fc82c454bf6b5943e3b

Observation 3c81927d-03e5-4188-a63f-84cfb6a3c2b0 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search Contrastive Localized Language-Image Pre-Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.876040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T15:18:17.427707Z digest=sha256:de3da46669d1b03453ccc19cafc5d796be634f58e146a0e36ede337e7678ea1e

Observation 2c58a7a9-7f40-4e70-a6dd-a1069d22512b · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search Contrastive Localized Language-Image Pre-Training

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:28.864002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T23:17:03.456746Z digest=sha256:ecaeb70db546463da862fd6ce285544e3c8c8c78a6e747cb24008c1ab54e3b13

Observation 7cb20dff-3047-436b-b11c-4a52cd1dafc0 · inbound

Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks cites this paper.

Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks Contrastive Localized Language-Image Pre-Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.790413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:16:21.252304Z digest=sha256:d34be87a6b3b20c46ba19fede868617d150492c7a3f31309737e52493eef69b9

Observation 29fc88c7-d200-4558-8e6a-94ca6844798f · inbound

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition cites this paper.

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition Contrastive Localized Language-Image Pre-Training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T11:31:04.191369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:31:04.191369Z digest=sha256:12b3f9374495093f50afcd807bbd14c34e167155608d7369527c5ad196671179