Pith. sign in

Paper Citation Record · LEDGER

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2505.24517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24517 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:06.804098Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:52:56.122056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f8decafb-89d9-422c-8e51-12b366787cdd · outbound

This paper cites Diffusion feedback helps CLIP see better.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Diffusion feedback helps CLIP see better

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:09.033380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:04.776421Z digest=sha256:a4c015dac4e3abb2cdb91028c6e0a100e79d1e5e42d4b3651e8d64ac6574d74e

Observation 2b0098d5-599b-4683-bec8-64439fc9930b · outbound

This paper cites GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:04.943789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:04.943789Z digest=sha256:545a68d39a49e4d0ffb1528f8e1507303b0ccb48a771733831de670ae72193a4

Observation b5cd51aa-8bbf-4623-a730-9047d014e94a · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP High-resolution image synthesis with latent diffusion models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:05.029316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:05.029316Z digest=sha256:2162da6552694be83b7dc40ea583d4d000972765da273034175d8d5bcd199bdc

Observation 9685119f-158e-4797-aca5-ffda49f9e336 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:08.853198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:05.126707Z digest=sha256:831d3ef85042f8ccd37d2ee7e092156130e98dfe4d48045244fe9ea9b76e8c34

Observation 5a750fcb-d2a8-4465-b1a5-6f6cf941383d · outbound

This paper cites CLIPScore: A reference-free evaluation metric for image captioning.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP CLIPScore: A reference-free evaluation metric for image captioning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:08.662957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:05.227208Z digest=sha256:d6b7d45a5f70b02e83cfd706897d14749aa8ce3d44bf7e0f682b39792ca32081

Observation a8ee72a1-d4c5-4330-a5c1-07e9e357b5b7 · outbound

This paper cites ImageNet: Alarge-scalehierarchical imagedatabase.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP ImageNet: Alarge-scalehierarchical imagedatabase

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:08.496592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:05.321714Z digest=sha256:c109de6914aa884a6a299a3cdd98542a9840bc5a32d17999c9a351bf4448dc2e

Observation 450a6e0c-ebc1-49a8-9eea-85a30eb2d1e7 · outbound

This paper cites Learning multiple layers of features from tiny images.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Learning multiple layers of features from tiny images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:05.445687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:05.445687Z digest=sha256:9d008e7229ee00148b8b0c3317913338e8359495fe2927d8d1a41b5779411363

Observation dac2d6b2-91ef-48ae-9261-cb508f2a8a55 · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:08.328385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:05.585534Z digest=sha256:b74482b8b82f7c69db0e34db550e633e7d56848116908c70d7dfcf73b7672633

Observation 9dc554cc-5143-40a2-bf09-1ffd4178d65c · outbound

This paper cites SUN database: Large-scale scene recognition from abbey to zoo.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP SUN database: Large-scale scene recognition from abbey to zoo

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:08.146328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:05.735773Z digest=sha256:f54472458c57eb89f08df2b7596bcf6d3ca8976bc64c180caa84f4a84070bb6d

Observation b058fb0d-93bf-4129-9bb9-8bedd0dd5634 · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Fine-Grained Visual Classification of Aircraft

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:05.810939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:05.810939Z digest=sha256:120eed5a1f4afd4ceee082738389abde80676e7e11efbe75167875d35ce8e3f9

Observation eb23a36c-aac5-4c53-ad97-4ac07fab9df9 · outbound

This paper cites 3D object representations for fine-grained categorization.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP 3D object representations for fine-grained categorization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:07.979623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:05.983789Z digest=sha256:550312b8344e93629beb0ff20b1af769e6c7de8d5b6916055fc78b85631b9442

Observation 04ae1e6c-089f-487c-bf8f-2cf426f8d10e · outbound

This paper cites Learning transferable visual models from natural language supervision.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Learning transferable visual models from natural language supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:07.864479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:06.099267Z digest=sha256:22b9ecf603663dd206414dfa4c1d4287f52ee86c379f9c1e2409d51f25d099b2

Observation db4a49ef-ec28-454e-8f5d-f190b68c98fc · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Reproducible scaling laws for contrastive language-image learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:06.207996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:06.207996Z digest=sha256:3b3fa35a8790f89910896092701ed20ced1ad262f014360ef9dc886c7f08a6ac

Observation 34cb4615-b0c3-4ee6-92f6-f98d34176b37 · outbound

This paper cites Kingma and Max Welling.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Kingma and Max Welling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:06.305565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:06.305565Z digest=sha256:e4467705edd9bebea64f87db5ecf42f7a30388e269ee6875f12dda9af873f5df

Observation 8113b8a4-38b8-4481-8a74-9671483e6906 · outbound

This paper cites Sigmoid loss for language image pre-training.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Sigmoid loss for language image pre-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:06.385555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:06.385555Z digest=sha256:034536a6b820aff0cad870fd676ec7d2217418a292a3248012b417d8027ce498

Observation 5c4840e2-e6fb-41ee-bd03-61833601c4ce · outbound

This paper cites an unresolved cited work.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:06.485464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:06.485464Z digest=sha256:5b825a15e38ba705c3266f354aa29168aa04e1c7d30a92f038e4496d352e3bee

Observation 24a119c1-ab39-42ba-90a8-cd0bcce3bff1 · outbound

This paper cites Microsoft COCO: Common objects in context.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Microsoft COCO: Common objects in context

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:07.592046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:06.580118Z digest=sha256:6f3c4b4b62f01524cc405c07f721ff5ca0e7d3aa71b8fa78cb1205036c197ed0

Observation a6cf2a6a-7620-4174-90ab-90c378d3418f · outbound

This paper cites Maskedautoencoders are scalable vision learners.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP Maskedautoencoders are scalable vision learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:07.396381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:06.692441Z digest=sha256:0ccc60f5f0b7a6354587e4239f77d29b9a26c7a93bc43505a7db1800686cac40

Observation bc77f57a-521b-4c72-8f61-d3d68647ff96 · outbound

This paper cites An empirical study of training self-supervised vision trans- formers.

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP An empirical study of training self-supervised vision trans- formers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:27:07.219444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:27:06.804098Z digest=sha256:1623ec01eba9bd6e240c942b3f3c138f8aa832f2c27fbee321dd42a6068ed1eb

Pith citing papers

Observation a18f549d-a4f2-4469-b88c-04a5501947f9 · inbound

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States cites this paper.

Contextualized Multimodal Lifelong Person Re-Identification in Hybrid Clothing States un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T16:52:56.122056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:52:56.122056Z digest=sha256:3ae7a174f0120d6dd183444429fca1edc1f507104e1c0d413bad6f318a4c5298