Pith. sign in

Paper Citation Record · LEDGER

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.22431.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22431 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:46:45.600246Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8687808d-6b90-469b-859c-9db72dba25ab · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.257026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.257026Z digest=sha256:3117e90f7873088c36fb317fa2cfcfd24cab8d2c4c7cef58ad587c0a321fc131

Observation fdb97316-3ec5-4106-9d2a-27cd89f67a64 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.727507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.264665Z digest=sha256:b0f88d7718f9a027e3bf99908932038743f187f4f719b698cc6ed2d26ce7353a

Observation e0755a2a-6891-4009-a98a-81960f9c5784 · outbound

This paper cites Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Internlm-xcomposer2: Mastering free-form text- image composition and comprehension in vision-language large model, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.697555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.275197Z digest=sha256:1f3b4fa8b768a618219ac9438575ade39b24a07225d27542b69622f4e6e978b4

Observation 0313582b-5d45-43cb-9d05-8a60e610ddbd · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.677311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.285435Z digest=sha256:2eb9d3aa9aa61c753ded93d8b264775b2e6eb4849dbdd3578a1ca8b3e5e859a2

Observation da9deaac-5bd3-4c4b-9d39-8fc13fc2bc11 · outbound

This paper cites Improving clip training with language rewrites.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improving clip training with language rewrites

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.656336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.293013Z digest=sha256:a7f91ef1dc0e444d31ce35553a991eb01e52117d73c212a49e79adc56dd2d00f

Observation 8494b625-6a4d-4b55-ab54-5a42840d6363 · outbound

This paper cites Data fil- tering networks, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Data fil- tering networks, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.638350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.306448Z digest=sha256:6c24a88c5a101de85c44bb8718359a2b5fde28cf7c0ad1e1e377d36c67276fc6

Observation 039d2dd5-6b27-4bc1-963c-aab84da425aa · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.616900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.313820Z digest=sha256:f074fb0fba365eca8cdde133258e9ce2de23eb986c905b2859e97ffd9085bd82

Observation 67bff9d6-c2f2-466e-b1f4-172f4223377b · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.588004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.325074Z digest=sha256:f8e99a633f6d9d7223cbf9722bf4bce72bf3cbd9f473c8b4de358dc60b63fb0e

Observation 08175ae1-0b0d-48a8-a002-ca6b7aa2389d · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Datacomp: In search of the next generation of multimodal datasets, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.567545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.332237Z digest=sha256:a87475993ba65f599ce1d2aa5e9f4b2ebb58e8a2dd15692d15a0cf28e16d0d2f

Observation 7262f8be-d511-49d4-9f5f-7e83a661c015 · outbound

This paper cites Classification done right for vision-language pre- training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Classification done right for vision-language pre- training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.547919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.339790Z digest=sha256:901d73236af31658bbd9f16923e0757191f1af66ffc19b6c0246534fcc11d5ea

Observation 911d6cae-6135-43ee-abee-333269196a7d · outbound

This paper cites GPT-4o System Card.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.348101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.348101Z digest=sha256:2edb52a6eb42f9ad2dae3e53a03b04c729f4dc6b525177f4b88f5eb40fc30aa6

Observation f54851e5-6dbb-4eb5-a7ed-24e0d60f14d3 · outbound

This paper cites Open- clip, 2021.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Open- clip, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.528791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.362242Z digest=sha256:fb490bb611bd59e4532cd1853cbe2fdf22fecbbcdaa02d14fa21f3bd3fa21e33

Observation f0fba9fa-74f1-4a7d-b16e-d53deb8a3785 · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions,.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Veclip: Improving clip training via visual-enriched captions,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.508293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.368122Z digest=sha256:9a7fb222213eaa0dabe55dc59c1b3e838ebaa699504cafd4db6377cea2631f73

Observation 372feb51-e920-4792-a9ab-c9d3bc5b0ee6 · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Seed-bench: Benchmarking multimodal llms with generative comprehension, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.486667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.375675Z digest=sha256:e15ca39d4cdfa69e093fb0e6080a3538a61a25d728ba51f2b71f3a1a034c4b35

Observation 77ed29c0-809d-412b-8fff-de2ece218112 · outbound

This paper cites Grounded language-image pre-training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Grounded language-image pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.457246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.387890Z digest=sha256:1cea9f18d26c2e76c15ec16103d3aecc004a80d2fe21bf708850b89f779682a6

Observation 3fe468b9-21f8-477f-980b-b93aadf25fa3 · outbound

This paper cites Clipa-v2: Scaling clip training with 81.17, 11.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Clipa-v2: Scaling clip training with 81.17, 11

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.428986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.395048Z digest=sha256:b9d4ecd9360b3f8705babb965b0d14114cbb21de476b351bbb107e6de52c8aed

Observation bca85c64-9114-4285-a012-c084a8a858ba · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.400816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.400816Z digest=sha256:eb1a1240273aa8a43ea54b5c147055f1ef7efd5af8669adbbdde4886b55f3725

Observation 12443c6b-bad7-4196-ab35-6c714c0a0a5b · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improved baselines with visual instruction tuning, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.410742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.415558Z digest=sha256:ef9bb1430c8b6a82de313c666d9b9dc2f6252115cfec30f36e3d7a6ab24be423

Observation cd5a54de-8e88-4196-9e92-ccfe55968d28 · outbound

This paper cites Visual instruction tuning, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Visual instruction tuning, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.420691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.420691Z digest=sha256:49f1f3333e8ac0539213e6b1903c0a67fdb4b54decfd5850fed5a83c5713dc14

Observation 60a38ff6-e757-4a52-91d5-6debf7111c78 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.376042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.427199Z digest=sha256:41d3a87b022483b718f0aa4fe63eedbc45ab3afcf5861a67febc3a3d243e33fe

Observation ab674b3e-c7a2-46c9-84cb-9726f7a5cfc4 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.432094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.432094Z digest=sha256:1f25bc1523ef0c7ea57dce5a26a371bec3a4f9824b5f3bb846cc0d9ca99077b1

Observation bae38e45-a8ab-44c0-866e-28cd93456e8a · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Slip: Self-supervision meets language-image pre- training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.331626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.440608Z digest=sha256:e44719649d5bf21ef4bd1a6828fec44c4b0e74f89bef7728d1956b19a4e7fbcf

Observation 51154b03-b426-4fd2-bc89-38e3bcd0ef82 · outbound

This paper cites Improving multimodal datasets with image captioning.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Improving multimodal datasets with image captioning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.308562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.447199Z digest=sha256:4e4349d52ef4b04c850187371f5d66b7d614e27b560e41ff1ab032560bca3373

Observation 1a6f594c-487c-42a3-be44-5dbf127470de · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.282647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.453397Z digest=sha256:bef087c82b504c48c6d46119c8b5ecd2fbd46928740e471983468c4f6d566082

Observation 21558c4b-2ba8-43ad-a5c6-19f6ce3b1c71 · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Denseclip: Language-guided dense prediction with context- aware prompting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.459856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.459856Z digest=sha256:bbfb6c463198f3e01ca3a91ad2211d0e92ade8667386180a99e0d7a039808654

Observation d1b1a601-d1dd-4e87-a7f7-aaddb8730b77 · outbound

This paper cites Fusecap: Leveraging large language mod- els for enriched fused image captions.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Fusecap: Leveraging large language mod- els for enriched fused image captions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.237591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.465588Z digest=sha256:8b22aea8f952a0a6ba0fa86e1d77118565d4fd4b6299e0441c1c26dc6f13da6c

Observation 972fd9d1-1252-4f40-af1c-da919df8a99f · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.476634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.476634Z digest=sha256:10c0b9ba1fc68c71af211ab62576728edc2f2d3a1f481935fadaab0432183708

Observation 3d5e3a2d-5cc3-4ce9-839f-03e51b64573c · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.214046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.483145Z digest=sha256:8a91a1920f0a301e843f2a0561440aefcb3df213277c07427f798bf5fc37b24d

Observation 41e5be2f-a1ec-40fe-b8be-6310f94a4b85 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.488144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.488144Z digest=sha256:6d595f9c1373051ee8a5b630e512ef80d11db232c4f49cb0cc079640325c4be1

Observation b64dc236-7bd4-4c47-b9e5-56161efecd1a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.494374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.494374Z digest=sha256:e9420d2b6f35ed64e2caee6d23b90d27cd075467f6af80024d04681f0ec313f6

Observation 2e124188-8a4c-472c-8dab-0138ecb71694 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.499725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.499725Z digest=sha256:ce0213c13e8421e86cf3f72f2c359f64c1ce12584782a7c8b8d44974abe0dd18

Observation 54248436-f71d-4947-a733-fae403771130 · outbound

This paper cites Locca: Vi- sual pretraining with location-aware captioners.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Locca: Vi- sual pretraining with location-aware captioners

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.184807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.505688Z digest=sha256:b6b747f29560a0365457fa53a38f498ab6b4b3ce73a18e30c9a5fb183448f274

Observation 1fbfab01-1955-4999-a42c-e42b7267011d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.512014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.512014Z digest=sha256:07e5ebff2b16f8d52480935c79c1a191a55b19754ca28481ccb5c503a5ec9ab4

Observation 2335abde-285a-4cac-8447-b1a2c9cbfa25 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.517124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.517124Z digest=sha256:d7fe6719f23a633269013a68f1bb3b8789b45282187c2676b5391fcc5b389cb2

Observation 04d2f84a-85a6-43b4-b30b-3156a9177e25 · outbound

This paper cites Demystifying CLIP Data.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Demystifying CLIP Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.524128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.524128Z digest=sha256:45f6c4aa972d1553ec6b3887d5a09a22112b4139ec252a219d8f286be1db04ac

Observation 8a8a5c90-6af3-4ba7-85d2-b63b449ccf57 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models, 2022.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Coca: Contrastive captioners are image-text foundation models, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.148854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.531912Z digest=sha256:f1419cb3a5f9b1a25f15acaa189d888bfc97babc69d203251f2eafa81ffb5614

Observation 85b5c121-2c3c-4e77-9e74-d93c9127a5d4 · outbound

This paper cites CapsFusion: Rethinking Image-Text Data at Scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models CapsFusion: Rethinking Image-Text Data at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.543788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.543788Z digest=sha256:9f110ea1acc25ae6cedca0ede0a9c42efd1616946b8d53c22701092109690f46

Observation c059c51c-33df-4d19-b842-cd199730cd5a · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.116925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.550364Z digest=sha256:31b4c15d28428d3cd11f90f42e315c555a3aaf858e878cc0d91a9b4a171d1884

Observation 8b7f51ce-0d25-44d5-9ed3-2353ba1dacc2 · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In International Conference on Learning Repre- sentations, 2023.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models When and why vision- language models behave like bags-of-words, and what to do about it? In International Conference on Learning Repre- sentations, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.091776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.558671Z digest=sha256:1e5822f3af2410fdfae25e4916db77818ead0c146fb2eaacef819f741d4e80f3

Observation 9669d846-da58-4512-9db6-0f32adf4edfc · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.568273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.568273Z digest=sha256:b2b34cea90ac2bc14aaf19470e9481123e27eaae959bc5aba18f5d56948df942

Observation 0a7d25ac-ca9c-4a7c-8755-c435c58b6959 · outbound

This paper cites Sigmoid loss for language image pre-training.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Sigmoid loss for language image pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.069538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.577179Z digest=sha256:70f0945c6e833c61cba9ba65aa1a19b6ac48cf191f7ae3b764a1c31365f766e1

Observation 8f631553-cbb2-41ca-b062-ad3de7d05214 · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Long-clip: Unlocking the long-text capability of clip

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.041210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.583471Z digest=sha256:afc3d4715770f2b87846b16b860cae4a44ea9d25892baa4fd29b57bd90597715

Observation 4adfeb5e-0eca-42d0-8e52-8e9c2088bd60 · outbound

This paper cites Glipv2: Unifying localiza- tion and vision-language understanding.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Glipv2: Unifying localiza- tion and vision-language understanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:46.010744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.588353Z digest=sha256:ad040797b1b251b794d53be6e47400826078ec03b3cbbb88f1a186a0797c8aaf

Observation d794a3e4-5197-44fc-963e-5adfd3f1bf62 · outbound

This paper cites Training setup and dataset scale.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Training setup and dataset scale

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:46:45.988293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.593767Z digest=sha256:9edc3f76e30798c19bcf8f8a8619057b8c703efaaeb5573b29f360b01a4a9d73

Observation 72ead4ff-f661-4f49-ac37-0df59345b06e · outbound

This paper cites Examples We present some examples from the acquired dataset.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models Examples We present some examples from the acquired dataset

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T11:46:45.961209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T11:46:45.600246Z digest=sha256:7be38d569a42f6b6f09ea96d5c7bec2b016ff4111029948ccca47b1fcc70545b

Pith citing papers

No inbound Pith citation observations are available.