Pith. sign in

Paper Citation Record · LEDGER

Xray-Visual Models: Scaling Vision models on Industry Scale Data

As of 14 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2602.16918.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.16918 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:26:00.495364Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b508cec-537d-41dd-8a86-794fba215770 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Symbolic Discovery of Optimization Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.349439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.349439Z digest=sha256:0cfcce2b82f265958adc2fbe64a6aeb277de0bb158968e16bbcd924d233c5f87

Observation 861bfd8e-0120-4558-bb20-82a55f20ac88 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.434888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.434888Z digest=sha256:10f30cd7472316c9d708101f6a082bba6e15aa8e5461dc1bc15711a0495a8d1a

Observation e0fc1917-471d-412f-96ca-d39244f146d2 · outbound

This paper cites Deconstructing Denoising Diffusion Models for Self-Supervised Learning.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Deconstructing Denoising Diffusion Models for Self-Supervised Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.514036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.514036Z digest=sha256:f3e8c89d1d0e576e2497fd08d84577d9292c140fe05e0f3b5828be3c8b8b8b45

Observation 5cc71124-0f00-41db-97c1-36ec5518baa5 · outbound

This paper cites Meta CLIP 2: A Worldwide Scaling Recipe.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Meta CLIP 2: A Worldwide Scaling Recipe

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.596814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.596814Z digest=sha256:83bc43936c429a0de6c3c2abadee143a94781783679b8c6362bee4636b59c3c1

Observation c1b7200c-2e27-405f-97c9-d69808dc5286 · outbound

This paper cites Vision Transformers Need Registers.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Vision Transformers Need Registers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.680370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.680370Z digest=sha256:762f48ea6e9ce5b1c7e5c553e933ce2db84634f4c1962619ef0798bef15d261b

Observation d7e408fc-18e5-4f02-b520-14d1d29dcb93 · outbound

This paper cites Improving CLIP Training with Language Rewrites.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Improving CLIP Training with Language Rewrites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.763379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.763379Z digest=sha256:d481b1be235095485646345263f13fb015adfd2f1573e7234d20a46d0bb1d991

Observation b594624c-d56b-4114-b2bf-95085c3e3573 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Xray-Visual Models: Scaling Vision models on Industry Scale Data SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.846646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.846646Z digest=sha256:c286b72640c27d39dfa915e3fb1d8515169e5099d45bf25e66706e66fcfedfe4

Observation f5676fdd-2261-48a0-a90d-caa06f22b53e · outbound

This paper cites The Llama 3 Herd of Models.

Xray-Visual Models: Scaling Vision models on Industry Scale Data The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.929501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.929501Z digest=sha256:79736218f794b5717ee32e3f12f761dce3fb7780020a5c8d057f4f42aab15d76

Observation 699fa30a-6d82-4273-ade3-3ced50b5e6ed · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Xray-Visual Models: Scaling Vision models on Industry Scale Data LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.074234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.074234Z digest=sha256:dd06d12dc165789eb3d4be52372b76b85f8a26d494742d9ffbca482c812a73c1

Observation 504cca27-18fb-4424-a97a-6feb7e8edbbc · outbound

This paper cites The Kinetics Human Action Video Dataset.

Xray-Visual Models: Scaling Vision models on Industry Scale Data The Kinetics Human Action Video Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.149510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.149510Z digest=sha256:0b2f4c192848c5d8fc5ca5027b40ad9386992cc93c03466d32aa519c471e26c8

Observation 01c48d91-fe66-4049-8a96-885d6737b0ab · outbound

This paper cites an unresolved cited work.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.314447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.314447Z digest=sha256:21546cc274dddc9db209860f8c2c21b0da32b69632a3063bfb4ad4048bd10b78

Observation 95340127-4d20-444d-9118-4f8ce75c5177 · outbound

This paper cites Exploring the Limits of Weakly Supervised Pretraining.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Exploring the Limits of Weakly Supervised Pretraining

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.397561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.397561Z digest=sha256:d1dcec3f100e49bd47b5f2c059d1135cf754450f95578151f4ee5ed0941d9b12

Observation f6fa7d81-3589-4fe0-a5cf-6eeadcac4002 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Xray-Visual Models: Scaling Vision models on Industry Scale Data DINOv2: Learning Robust Visual Features without Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.478503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.478503Z digest=sha256:ab1ac4640d9b93cb27723a13f67958d7bd976502a12e28a17c80762fb9f80f66

Observation 696c96a6-75cc-40bf-b530-93bb54c9f38a · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

Xray-Visual Models: Scaling Vision models on Industry Scale Data LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.535987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.535987Z digest=sha256:cfa0b5990365210853ec9d579ca088751014b66eb6abe7b056ad7819303c3e57

Observation 3dbfe0ee-661d-4a8a-9b15-41a9782eab60 · outbound

This paper cites Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Adafactor: Adaptive Learning Rates with Sublinear Memory Cost

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.611339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.611339Z digest=sha256:beff108ba8daf283a745299e3e8c78b4fbc1bf829e2a09b0af9e32a519acdc3c

Observation a4cc25f9-992f-4ba6-9599-89d16fb03e88 · outbound

This paper cites DINOv3.

Xray-Visual Models: Scaling Vision models on Industry Scale Data DINOv3

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.691930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.691930Z digest=sha256:ed1ecb8ec50d8ac60ac684e7308e624e6c9044e970d5a0c308be6d01ae435b86

Observation b55b2978-a860-4702-a0b5-39b37c7870e3 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Xray-Visual Models: Scaling Vision models on Industry Scale Data RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.752258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.752258Z digest=sha256:82bf5caed9f4e495fd8d797baf62a6daa4674c230b196fa6c47ae01c8b557813

Observation 9e0582b5-d228-4703-97d3-dc1607de51c5 · outbound

This paper cites Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Pooling And Attention: What Are Effective Designs For LLM-Based Embedding Models?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.821475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.821475Z digest=sha256:fad02a8872d4f5d0b19104d7c6eee5aeb0190bdb81e2e8461fe958367d7ce0b7

Observation bbc80d26-e53f-4a8a-aa16-25f31f6abe37 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Xray-Visual Models: Scaling Vision models on Industry Scale Data LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.878103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.878103Z digest=sha256:8a904996a870a56f0e049bcb43223a92307a19e420248e578b011410f8ed7e5c

Observation db195f75-9445-4edb-8d16-bd3c9152f423 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Xray-Visual Models: Scaling Vision models on Industry Scale Data SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.954000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.954000Z digest=sha256:7bb4b222526d6fc0ac798e00fe33af3ea4394d884c78e0f723fd503f8cee9b3f

Observation a03cbcfc-fb3c-4eed-b62c-572431ca23dc · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Representation Learning with Contrastive Predictive Coding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:00.039076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:26:00.039076Z digest=sha256:b509ef0a44e630f7c53bb3c50fdb02ea028a311e3f20835d6c303be605ed5272

Observation 0966b5a3-9fd4-4390-aa95-823c9abd1975 · outbound

This paper cites Demystifying CLIP Data.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Demystifying CLIP Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:00.202547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:26:00.202547Z digest=sha256:cdfd8b2663aaaf82c7bcb70d5b411aa58949ccb5912d1a4d5bc13a1819afaf9b

Observation 1a317e63-dab9-481c-8136-3d3b0d18d5e5 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5288–5296,.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Msr-vtt: A large video description dataset for bridging video and language.2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5288–5296,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:00.317594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:26:00.317594Z digest=sha256:797b95944004244fac13a10002b2ee89e59f7a29da31b761050fada67a5ebf9f

Observation 4b902530-fcc4-4c89-8711-732db56544e2 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Deep Residual Learning for Image Recognition

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.015365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.015365Z digest=sha256:f63542d140588f75ac4924390abe85c83bda731d5e3e806b678d4c48ac63b3ab

Observation e4a64b7c-4877-4d2d-9b1b-2d934359eb4f · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Xray-Visual Models: Scaling Vision models on Industry Scale Data InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:00.495364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:26:00.495364Z digest=sha256:c8b794dcce310bbcee6d5f37e8c22f6882c0e53b22fe2fff4c07f2c7461a18d1

Observation f5522b7c-7839-4be4-8dce-a74a5544fd79 · outbound

This paper cites Hong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang, Marcin Eichner, Keen You, Meng Cao, Bowen Zhang, Yinfei Yang, and Zhe Gan.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Hong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang, Marcin Eichner, Keen You, Meng Cao, Bowen Zhang, Yinfei Yang, and Zhe Gan

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.269295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.269295Z digest=sha256:780b2b3a1a41aa6530cd4a3795d9fa70c18941130e9a8c9f6943057124c7c221

Observation 1e7d7fe6-1d9e-4aab-9d44-46a5177bd1ac · outbound

This paper cites LocCa: Visual Pretraining with Location-aware Captioners.

Xray-Visual Models: Scaling Vision models on Industry Scale Data LocCa: Visual Pretraining with Location-aware Captioners

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T22:26:00.121160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:26:00.121160Z digest=sha256:ffb351370a95d983e6332c490aad32d4449d2e961ea7d376e6f1cdb8b2aa9945

Observation 77d81c88-65f5-4b27-a17a-116f7dfc034f · outbound

This paper cites Language Models are Few-Shot Learners.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Language Models are Few-Shot Learners

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.087918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.087918Z digest=sha256:9d4c2963a307025754f2d13b29abe127698948ae3d74e402084ddde6e09cd307

Observation eebaee77-cf9e-4e9c-b509-705411c184e5 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Quo vadis, action recognition? a new model and the kinetics dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.185292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.185292Z digest=sha256:209f10edc5248c91716e150331f01bee72d6dc0431291e391563db85880beabf

Observation bc6b390b-0fa7-4974-ba07-ac700ebf8f6d · outbound

This paper cites Improving fine-grained understanding in image-text pre-training.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Improving fine-grained understanding in image-text pre-training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:57.884381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:57.884381Z digest=sha256:8f49a32729c6a75c385c5bb7df520031575b6c5d2990f044cc0eda16d16cfb3a

Observation 67e131ad-0a4e-46a9-9c23-526340582a32 · outbound

This paper cites Poggio, and Thomas Serre.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Poggio, and Thomas Serre

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:59.231880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:59.231880Z digest=sha256:a94aebe80878a1d40333b790dc20a76f79f98c15cf23323d7baf77e12fd5e0f7

Observation 3a5f23d3-fca3-48fc-bbbe-f97a1817f0e2 · outbound

This paper cites Token Merging: Your ViT But Faster.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Token Merging: Your ViT But Faster

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:57.964997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:57.964997Z digest=sha256:9b1a62c38b94e4e0e68154e6488655032e3b614f290a8dcd9720ecb076cfae51

Observation 8ca2d4af-bf68-4a7a-b8bf-f51fae9b828e · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

Xray-Visual Models: Scaling Vision models on Industry Scale Data V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:57.827119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:57.827119Z digest=sha256:1529c4ab0f294bf264ddb8d943e549f35d9a68e2fca807838fd45b976dd7ce4a

Pith citing papers

No inbound Pith citation observations are available.