Pith. sign in

Paper Citation Record · LEDGER

Meta CLIP 2: A Worldwide Scaling Recipe

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 22 inbound Pith citation observations for arXiv:2507.22062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22062 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:08:23.842628Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:50.676419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.777855Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d034e13-e2da-40ca-b488-e82fea865f11 · outbound

This paper cites Towards Zero-shot Cross-lingual Image Retrieval.

Meta CLIP 2: A Worldwide Scaling Recipe Towards Zero-shot Cross-lingual Image Retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:21.963849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:21.963849Z digest=sha256:f6133a988bb2ba800448ca4cbd3e76369af5c59f6ac14306c7dbeb687ac4aa18

Observation 5c662cf9-bafa-4fd2-9608-654a664a0bd5 · outbound

This paper cites an unresolved cited work.

Meta CLIP 2: A Worldwide Scaling Recipe Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T12:08:24.446114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:08:23.842628Z digest=sha256:9046d91820b3e8f723d5de58da9dcbcd1ab48c532bd80e5c973113d7869e7f5f

Observation ff29788c-20c9-4cd0-b73c-f8bff230d0e4 · outbound

This paper cites Learning word vectors for 157 languages.

Meta CLIP 2: A Worldwide Scaling Recipe Learning word vectors for 157 languages

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:25.102057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:08:22.555512Z digest=sha256:ef6f6983f5024c8fb70a848914f51136d9d6898c1f103a544ee4e0c4a1dd463f

Observation 7649df2c-aa6d-447c-96c1-0a43d64893d2 · outbound

This paper cites Graph-RISE: Graph-Regularized Image Semantic Embedding.

Meta CLIP 2: A Worldwide Scaling Recipe Graph-RISE: Graph-Regularized Image Semantic Embedding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.817078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.817078Z digest=sha256:b637712e9a2424db69a8f4fd86226370524e375828a19b6a3a716450d4cfbb0c

Observation 75adb5fa-a286-4cdc-b65e-24144d4bfef7 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.International journal of computer vision , 128(7):1956–1981,.

Meta CLIP 2: A Worldwide Scaling Recipe The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.International journal of computer vision , 128(7):1956–1981,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:24.913110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:08:22.884214Z digest=sha256:28c513ffa097c7562524fccd7a582a4b9d65736f8bfa1f7d0d066dd7b94d4b1b

Observation d1d3d36e-405a-4426-bff0-9baa7f1a9f13 · outbound

This paper cites XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models.

Meta CLIP 2: A Worldwide Scaling Recipe XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.953395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.953395Z digest=sha256:be8eb9fa24801f4d044a421eeb85c7250629faf73ee7529b4114b2915bb7857a

Observation 026b57e2-414c-42b3-ac8e-6a91ed6ce1f7 · outbound

This paper cites SLIP: Self-supervision meets Language-Image Pre-training.

Meta CLIP 2: A Worldwide Scaling Recipe SLIP: Self-supervision meets Language-Image Pre-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.020444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.020444Z digest=sha256:fec2fdb26beea9879d2f0c087ff6407df35bb8b405235b02b9275c2d142d83f3

Observation 31426624-87e8-4e6a-8b42-cb9d16d1b56b · outbound

This paper cites CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages.

Meta CLIP 2: A Worldwide Scaling Recipe CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.119878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.119878Z digest=sha256:b311d006f02d4a6458e931da0d955cf0c2664dfb4684b7042b44a43a57cd627a

Observation 281372e5-4acb-46cf-b6d6-69c70178464d · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Meta CLIP 2: A Worldwide Scaling Recipe LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.191317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.191317Z digest=sha256:ef8a29051860f377ebd8f646a20295afdef4e716d700e1a9533f438862e39960

Observation 2ecc4e40-0b38-4074-ad82-498fc78157c8 · outbound

This paper cites No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World.

Meta CLIP 2: A Worldwide Scaling Recipe No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.261268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.261268Z digest=sha256:297eae2d64d01026d675903801323bdc226c03521dcc29fd038c2474dcabb057

Observation 98d8d6e3-ba55-4421-9c91-ad6d49c8643f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Meta CLIP 2: A Worldwide Scaling Recipe Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.326200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.326200Z digest=sha256:a7db352adcd00ab7f5191ac68a662783ce76087db981a247214b1559dd3a8514

Observation fbeabe99-1b29-4ede-8dd0-70a37045b228 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Meta CLIP 2: A Worldwide Scaling Recipe Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.350759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.350759Z digest=sha256:87854a7dac4a30650ac6e05ecb10688f594eb1288e120cd4edc3a028c67b553e

Observation 8f4f2302-1313-49f0-9a4e-de621667b2b9 · outbound

This paper cites Crossmodal-3600: A massively multilingual multimodal evaluation dataset.

Meta CLIP 2: A Worldwide Scaling Recipe Crossmodal-3600: A massively multilingual multimodal evaluation dataset

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:24.745850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:08:23.369919Z digest=sha256:d28317be8cb3bf11db324ec84c2e5ded98a922128d25380c915dd80f9daf1ec9

Observation f0a212a6-4dde-4bc2-b561-d028cfc6cc4e · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Meta CLIP 2: A Worldwide Scaling Recipe Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.444125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.444125Z digest=sha256:f2f3f1c92eeca49b2f1b9f005d5a6fcbe005209f8036d9f6d7ef9a6a69bc42c0

Observation bfcc85ce-9b18-495c-a5e2-93236e1e3618 · outbound

This paper cites NLLB-CLIP -- train performant multilingual image retrieval model on a budget.

Meta CLIP 2: A Worldwide Scaling Recipe NLLB-CLIP -- train performant multilingual image retrieval model on a budget

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.507484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.507484Z digest=sha256:83841762b5ed8ad8a5126e70d90f1d9e2faec6c803111cd2e2be51143d399d9c

Observation a9d2dc30-1e99-4169-bad4-5885be5b972f · outbound

This paper cites Scaling Pre-training to One Hundred Billion Data for Vision Language Models.

Meta CLIP 2: A Worldwide Scaling Recipe Scaling Pre-training to One Hundred Billion Data for Vision Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.606804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.606804Z digest=sha256:6fd4626848ec2cce6fb34d538d06001d166596fe93adc1f4bd0e3f3bec53333d

Observation 39208fe8-736d-474d-8ff4-cdd9ef151fa0 · outbound

This paper cites Hu Xu, Saining Xie, Xiaoqing Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer.

Meta CLIP 2: A Worldwide Scaling Recipe Hu Xu, Saining Xie, Xiaoqing Tan, Po-Yao Huang, Russell Howes, Vasu Sharma, Shang-Wen Li, Gargi Ghosh, Luke Zettlemoyer, and Christoph Feichtenhofer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:24.592885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:08:23.693138Z digest=sha256:fae1bcce80eaf89268c97a9802afcf1cee5bb803bda3bae25f14e0e1ce97af5f

Observation c7322498-8406-416d-bca1-338d1589208f · outbound

This paper cites mT5: A massively multilingual pre-trained text-to-text transformer.

Meta CLIP 2: A Worldwide Scaling Recipe mT5: A massively multilingual pre-trained text-to-text transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.768605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.768605Z digest=sha256:2d556cdd06c8bbf189a6e7b591de576206e6cb7b3631c0cd588afa6016852a12

Observation d444a0ac-efec-4fa8-bf70-ba13f8d9b634 · outbound

This paper cites Scaling Language-Free Visual Representation Learning.

Meta CLIP 2: A Worldwide Scaling Recipe Scaling Language-Free Visual Representation Learning

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.283340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.283340Z digest=sha256:81235c3d88cc046284fae45c0326ff68075e405cceb3d6d6f3442621045cb4e6

Observation 1b915d3e-10f0-470c-b0f5-1145895b90ee · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

Meta CLIP 2: A Worldwide Scaling Recipe Unsupervised Cross-lingual Representation Learning at Scale

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.145470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.145470Z digest=sha256:bfd7cbcca1a648880933d0ca540ffac07956cb0d13eb5b279eb8c12df3bd2bad

Observation 92b222e6-291a-482a-a18f-034b2357914e · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Meta CLIP 2: A Worldwide Scaling Recipe Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.297514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.297514Z digest=sha256:69342a857208340683a79db2f278cbf433ba92ced5fa81633901f9b15a538391

Observation f75b13bc-7463-4314-9967-2e87e14040ed · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Meta CLIP 2: A Worldwide Scaling Recipe Distilling the Knowledge in a Neural Network

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.646735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.646735Z digest=sha256:2ae6402f057a5f3047525b1e76649ae32f442939fbd1f7678d26fc7ab944ea77

Observation a7c59474-99bf-47ce-b19a-0831ef121dd2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Meta CLIP 2: A Worldwide Scaling Recipe Imagenet: A large-scale hierarchical image database

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:08:25.248608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:08:22.212731Z digest=sha256:3ade3c2fb439557b7d4f0d1b99f2edd257b5d61d07fa88351cc9d7aca426fe6a

Observation 9666c542-7268-46b5-be0f-26fd43753bdd · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Meta CLIP 2: A Worldwide Scaling Recipe Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.071806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.071806Z digest=sha256:c1c1e06687add8d2c798f478ac1ec53412f607d7b6fef8c16c858f3e4745b54f

Observation 86a79199-d5fd-4a11-a37a-b19ab22ad20e · outbound

This paper cites If you use this software, please cite it as below.

Meta CLIP 2: A Worldwide Scaling Recipe If you use this software, please cite it as below

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.740432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.740432Z digest=sha256:e445c9600da22c2c4f25417033f8df046e86bb72220c2a58f79cf720cc8b11da

Observation 249240a4-2ef1-4e62-bc20-b991b752b8c4 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Meta CLIP 2: A Worldwide Scaling Recipe SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:23.416279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:23.416279Z digest=sha256:50e6cf88449ddc7a72bb46511eede51cc2e003ce6baa90ec2f3bda49fb482e08

Observation 7f6ab0a0-fc9a-4a6a-95b0-28bbcd94b918 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Meta CLIP 2: A Worldwide Scaling Recipe Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:21.997481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:21.997481Z digest=sha256:4ef5b93dca59f907039bbce6e995e0bf8480f5fa37b740d09d9221365e02f9fc

Observation ad84ca00-79a6-412f-85d1-f5c3e8ab55c5 · outbound

This paper cites The Llama 3 Herd of Models.

Meta CLIP 2: A Worldwide Scaling Recipe The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.487941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.487941Z digest=sha256:4f41fb26ff8b11b29cea6e6d42477e1656999e16c360f2b90e6a12cc7d8b6989

Observation d8b31b26-ff84-4152-815b-05776a968cd8 · outbound

This paper cites Data Filtering Networks.

Meta CLIP 2: A Worldwide Scaling Recipe Data Filtering Networks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:22.412069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:22.412069Z digest=sha256:7c9ab8447c1225c12d6892b079f681c7ab1d7348f182ec318a0dffd896191cac

Pith citing papers

Observation 5332ff76-ab9f-4b47-a996-dc1aa47a076a · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Meta CLIP 2: A Worldwide Scaling Recipe

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.469427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:36434ecde73920b850225c5181e229f4dcd987e9d27dc6f0047ccacd51410e67

Observation 5ce06e82-7188-48f2-8e50-03db7c0ebcbe · inbound

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval cites this paper.

GRAPE: Let GRPO Supervise Query Rewriting by Ranking for Retrieval Meta CLIP 2: A Worldwide Scaling Recipe

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:21:21.357576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T12:18:06.259724Z digest=sha256:dba94cde70e663e96a5941b5fe85015176eb0373a6ffc4edc9adf9cb79f4d9ed

Observation cc627568-7ec6-4a7e-a864-c6d92a5ffca6 · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Meta CLIP 2: A Worldwide Scaling Recipe

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:42.782278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:42.782278Z digest=sha256:618f325435ba5210d770b8b0e8fcf123af9c2815ee0c66772d0fa6695f4247c1

Observation 56e06a15-ef53-492f-a6c2-85132b4de802 · inbound

PowerCLIP: Powerset Alignment for Contrastive Pre-Training cites this paper.

PowerCLIP: Powerset Alignment for Contrastive Pre-Training Meta CLIP 2: A Worldwide Scaling Recipe

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:59:03.966572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T04:57:44.794342Z digest=sha256:8e612ba4cad3e0ebfffb11e6919b2cefd0c744a66ca108d1bc9aac405ac04cb7

Observation 97df5930-a249-47b2-bfd1-82bba7245d8d · inbound

Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models cites this paper.

Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models Meta CLIP 2: A Worldwide Scaling Recipe

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.262899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:40:09.385451Z digest=sha256:493065f33a4b0104c50ea84486dd4ca4cf3efabe1e54465e4ddf737ff964aeff

Observation 5cc71124-0f00-41db-97c1-36ec5518baa5 · inbound

Xray-Visual Models: Scaling Vision models on Industry Scale Data cites this paper.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Meta CLIP 2: A Worldwide Scaling Recipe

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.596814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.596814Z digest=sha256:b5a0e9442b25637c8e0245f57966d0f72dd144f5751d06502840f607f7c9a4dd

Observation f1a22a00-6cc0-4285-8f90-441f399b94e6 · inbound

Peel neighborhoods cites this paper.

Peel neighborhoods Meta CLIP 2: A Worldwide Scaling Recipe

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T20:08:04.785209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:08:04.785209Z digest=sha256:37da10ea80ddc713c75a34fa7cff9a9c2c6cc62c94c47cfb64e65b729728f9f6

Observation d49a169b-8bc2-470a-895e-0ce8b4fcf39e · inbound

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models cites this paper.

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models Meta CLIP 2: A Worldwide Scaling Recipe

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.021625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:19:04.468630Z digest=sha256:e1e5aaf79bd3cbd842f369fbd59afb9c9a7c4a6642f83e9bafac8ce057f3292f

Observation 0bda72bd-3ab8-4aa5-8ca3-ab1823d882fc · inbound

HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild cites this paper.

HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:58:09.042956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T18:53:14.522561Z digest=sha256:efaab43d08e523ec73c299448050bdf81e2f4512716888feedc984f8f2263b3d

Observation 113b0666-aff2-487b-b8b9-9c5224cb1394 · inbound

LOGER: Local--Global Ensemble for Robust Deepfake Detection in the Wild cites this paper.

LOGER: Local--Global Ensemble for Robust Deepfake Detection in the Wild Meta CLIP 2: A Worldwide Scaling Recipe

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:38:07.133714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T18:37:11.529615Z digest=sha256:78e0011a2f0d01faa6dda8c37978345f89d3ec23f13f67f4610335390239bae7

Observation c9e66fd2-8d9a-4934-864f-971ae008fe16 · inbound

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving cites this paper.

SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving Meta CLIP 2: A Worldwide Scaling Recipe

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:01.620906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:52:37.723921Z digest=sha256:83950bcd30877ccb33f0d49eaf9e07a417ea321c480cf14f90443ea1b3a3c1dd

Observation 2842cd12-81a7-4d8a-944c-a8efe8790ae3 · inbound

Boosting Robust AIGI Detection with LoRA-based Pairwise Training cites this paper.

Boosting Robust AIGI Detection with LoRA-based Pairwise Training Meta CLIP 2: A Worldwide Scaling Recipe

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:01.868275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:56:11.099966Z digest=sha256:888683b61dac7e35a2a104cbeaa95b2bd5391fa61f365696c849168cdb47bd27

Observation a63c8dcf-d310-452d-8cad-dc0de84359df · inbound

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding cites this paper.

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding Meta CLIP 2: A Worldwide Scaling Recipe

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:05.030263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:24:57.169737Z digest=sha256:c3a4b69d1ef94e78414f50e1c1666dada2be282caed397c6ae309669b81ddc89

Observation 985bcb05-b485-409e-95e4-6b89a8d86d05 · inbound

Sapiens2 cites this paper.

Sapiens2 Meta CLIP 2: A Worldwide Scaling Recipe

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:06.956239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T21:59:43.755956Z digest=sha256:8685a4a9eab862d678a506a8b9eb1cc0e3c314f386de44b169187fc51bd9f764

Observation f69a7296-e811-4d1f-a3a1-280deaa44782 · inbound

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries cites this paper.

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries Meta CLIP 2: A Worldwide Scaling Recipe

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.348924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:22:14.870351Z digest=sha256:c0bd2046e07382675120f5de81b8bd480fe73cec2ba28c5fc9f04850394b4754

Observation 3eaeb297-59c8-4cec-ab70-77fe8fb0f498 · inbound

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis cites this paper.

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:18:00.153547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:10:36.000486Z digest=sha256:6d444eb23634b3f2462ceecfc94f79b55428e36ba1fbb5555875b9bffb531a63

Observation 2d0bd3ec-2eeb-4629-abfb-16f5a8519ac1 · inbound

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain cites this paper.

Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:46.323860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:40:00.643198Z digest=sha256:6d44ddb32996b4fde1836e22267b9e5719cf54ce49856ada55c727c16ed51947

Observation e3f4c34d-e547-49c8-bfe1-d21041147b67 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Meta CLIP 2: A Worldwide Scaling Recipe

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.779287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:06e8f7e384a3215bf5cb64844f74ca7418fb446162ac6bca455f81b5c2263d3e

Observation 9b619a9b-4c31-4cf0-bcb1-e3e75745f5fa · inbound

AdaBoosting Text Prompts for Vision-Language Models cites this paper.

AdaBoosting Text Prompts for Vision-Language Models Meta CLIP 2: A Worldwide Scaling Recipe

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:08.304554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T16:24:03.231930Z digest=sha256:21f7bf7cb9df7b3ac851d6b9ef3295f0a8d465515ce5d5c7194a3abb5ec6f096

Observation 85cbf92b-cce4-4563-9cfe-942b987fdf07 · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Meta CLIP 2: A Worldwide Scaling Recipe

Reference 106

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.029211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.029211Z digest=sha256:1050e2202fdfc73f9d7ac68e085fdb9a8a9a1a03bf07964d250226adb5f3733f

Observation 6d0e6292-82af-4548-95f5-d21db5e3b9d6 · inbound

Fine-Grained Food Image Understanding via Target-Aware Data Alignment cites this paper.

Fine-Grained Food Image Understanding via Target-Aware Data Alignment Meta CLIP 2: A Worldwide Scaling Recipe

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:32.988872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:32:32.988872Z digest=sha256:9d7b5844016b59700a6f63730a87b7dec5e22e9a5c3e5ff50d092a739065c362

Observation 5c7b89bf-88f1-4b03-9f7e-bfa49d0f73af · inbound

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning cites this paper.

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning Meta CLIP 2: A Worldwide Scaling Recipe

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:50.676419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:39:50.676419Z digest=sha256:e33243a71c75416f1ce97637106af8288de73ecd4ba33a0c1e23fe02958196c5