Pith. sign in

Paper Citation Record · LEDGER

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling

As of 17 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2604.13561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.13561 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:49.805446Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b13967f2-5868-43ee-8a58-ccaffc36f58d · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.833511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:cb0e6d304d4faec2c72728da6745137d8969a950170d4bd31e2230ba5279e572

Observation 4702c512-9832-4f27-b6be-c6929d300968 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Scaling up visual and vision-language representation learning with noisy text supervision

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.859678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:4ffc942ac7eb7421bebd29dfb666435b23b8c2ad98c329a1c905cf4a0cdd3b1c

Observation 62bbe677-aa85-49c9-ba86-e96b53c2ed73 · outbound

This paper cites Contrastive learning of medical visual representations from paired images and text.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Contrastive learning of medical visual representations from paired images and text

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.862502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:6427bd411f9fd0db941603f66e58e485310f32cb4bd9c3b8efa30248caf2d58d

Observation 5919836b-2db9-469c-9008-55ad67265a19 · outbound

This paper cites GLoRIA: A multimodal global-local representationlearningframeworkforlabel-efficientmedicalimagerecognition.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling GLoRIA: A multimodal global-local representationlearningframeworkforlabel-efficientmedicalimagerecognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.842510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:1d8436edce3b18b1c026e3aa007f4a3f5807af61e40a525b85f3038fa1ee920a

Observation 3bec40d3-9901-4fd6-ad5d-d14cef151f45 · outbound

This paper cites Mak- ing the most of text semantics to improve biomedical vision-language processing.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Mak- ing the most of text semantics to improve biomedical vision-language processing

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.839704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:9d34f8c63b0cf4fa8333f6211569703471b81f032f24e70ed14008e68803c8c9

Observation b5228490-9216-4919-a15c-3e5cd2f28d64 · outbound

This paper cites Merlin: A vision language foundation model for 3D computed tomography.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Merlin: A vision language foundation model for 3D computed tomography

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.830758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:4eccbd8b662138917b2350f779fcc393629159acc6898c25d5d56af75b320e5b

Observation eaba4fae-42d1-474f-91fa-99688e14c803 · outbound

This paper cites Balanced con- trastive learning for long-tailed visual recognition.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Balanced con- trastive learning for long-tailed visual recognition

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.868472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:e05b307b8d0fb35518e50de9b53d654ff991312b56264fb84bb88e0f70b6f02e

Observation d179ac1d-c08f-474e-8301-da5e653fe23d · outbound

This paper cites Parametric contrastive learning.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Parametric contrastive learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.827976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:5a4eb9d55acf462eec8bb9b2ce6a703500ffff2e5742a805e6635b9127618ebd

Observation c755be96-a447-43c3-9796-ca606bf9160b · outbound

This paper cites Contrastive learning with hard negative samples.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Contrastive learning with hard negative samples

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.853437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:90627324c913081c9cde025f6d4a08d4c7ca5569c046609d22ad4887440ebd6b

Observation a7eac4ac-7a45-47cb-9614-b6b3c58c4f43 · outbound

This paper cites Hard negative mixing for contrastive learning.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Hard negative mixing for contrastive learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:46:50.549842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:7a1c3f764bbf637c3a03d725937eaaa794c03361203b6237643d6a5fea927a1b

Observation 306ea2c7-563f-488e-a547-aafa8952523b · outbound

This paper cites A simple framework for contrastive learning of visual representations.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling A simple framework for contrastive learning of visual representations

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.836605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:a4398ccaa47a0241196920734a9ea4e35bdb1c65326907881e2aa7d3b0bb2da9

Observation b06e1f50-b68a-47ef-b8e8-d8f45e12f4b9 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Momentum contrast for unsupervised visual representation learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.825134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:92a7194cb6c95ec8a72ecade91546bdb1013c5f6d1e1720407cd0a0284e58d1b

Observation 8a00a620-201d-4567-97fb-c101ee241909 · outbound

This paper cites Decoupled contrastive learning.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Decoupled contrastive learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.847955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:a4fe8e83ff07c1da03f5baebfeecba89aea56da9b5ab06cb7b3d6063a0411df8

Observation 39c01186-d643-4cf3-a9b3-5253b18e7fb1 · outbound

This paper cites Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.850659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:e8cf5f86ade4a9c7b1d124fa60b2ee14d9b759b97906f447fc8d4434f6132180

Observation 1c0b6af8-f654-4ad3-9399-07a05f14ba47 · outbound

This paper cites Generating CT images from free-form text reports using GANs.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Generating CT images from free-form text reports using GANs

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.856612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:d1ed1a988601bacee7c8e56cb7ac8e0de7b3c665c16cdd100f9f6cae96756389

Observation 51e8502a-2a4f-497d-ba36-fffd7d1982a1 · outbound

This paper cites Reproducible scaling laws for contrastive language– image learning.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Reproducible scaling laws for contrastive language– image learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.865335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:79c111ad4a66441aab733b08519b30b39cdfe7411685aafe430de39ca7f56938

Observation ac39b4b7-dbab-49bb-9924-c227e5ffc80a · outbound

This paper cites Deep residual learning for image recognition.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Deep residual learning for image recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T19:51:50.845287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:9ced307d2be40fe444e28cfc69661bfd1f8174aace7e2d824b288e17ee918132

Observation 5ef4fcc6-6eae-4bb7-b927-44b2c4be2020 · outbound

This paper cites Clinical-Longformer and Clinical-BigBird: Transformers for long clinical sequences.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Clinical-Longformer and Clinical-BigBird: Transformers for long clinical sequences

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.806088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:5da01593a02aae60dc1d63a3211274af7fbef27a06ce5b59bffe48f3692b4eb5

Observation 872b9791-e43e-40dc-9145-a808c37fa2c1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling Representation Learning with Contrastive Predictive Coding

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.800372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:49.805446Z digest=sha256:5eb1b08e390e7878277037b13716d2b6dc4b920254e661fd9df8d78fbdd8d5f8

Pith citing papers

No inbound Pith citation observations are available.