Pith. sign in

Paper Citation Record · LEDGER

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2102.08981.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2102.08981 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:38.530639Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T02:48:44.958361Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5e6d4d93-75e6-459b-9683-df9a65d5b809 · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:26.919850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:e54a78c01ece696df4b96da980e8b26f1cd11318912cfb3b67ee24663c962f10

Observation 4f631775-823a-447f-ad2f-eb68bc24f782 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:44.961776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:465fd12d00bd931d1c2aff9c3ea96dac5cb4fc1f730ad80edd7fe5b09971413a

Observation 4b1d9c74-8136-437f-94e8-736a37685071 · inbound

Native Segmentation Vision Transformers cites this paper.

Native Segmentation Vision Transformers Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.530639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.530639Z digest=sha256:591b46f32242d7c4e03986590f5c61d2aafc5d43d2ba42d2b992ea28ed47e18c

Observation 24bc5e7d-237e-4aea-a8b8-1b81580e35b4 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:59.063835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:59.063835Z digest=sha256:4c4f589f7f5d14b2e69f48b3da8c074e85c96203e29670f4cdec8619a8e372e1

Observation 2ad43a37-cd7d-49f7-9fb6-bd9a39efa155 · inbound

Entity Image and Mixed-Modal Image Retrieval Datasets cites this paper.

Entity Image and Mixed-Modal Image Retrieval Datasets Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:45.661946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:45.661946Z digest=sha256:0919ff508c6a0745872e7b3231e329de44f5b76b06487ea29d484dc8a5cf5273

Observation 8dc5c7c3-b4f4-4086-9b32-c5b43d10ab15 · inbound

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation cites this paper.

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:59.977186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:59.977186Z digest=sha256:96655076e426b1c2fffe817a4d731ce83f4f6a0d8ca32765dbbaba1a5a27113f

Observation 896a3b3e-2969-4853-bef4-6ce1b59f5d62 · inbound

Info-Coevolution: An Efficient Framework for Data Model Coevolution cites this paper.

Info-Coevolution: An Efficient Framework for Data Model Coevolution Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:32.928691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:28:32.928691Z digest=sha256:570e7ef17d75244136a6799d546f67599affe9d683767071e15c9997914bcb41

Observation 56413fac-9e7e-4d3f-84a1-11cbf214d70e · inbound

SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation cites this paper.

SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:00.098365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:00.098365Z digest=sha256:6dc72fa3a2aeb495af6e77e0c2fa90c2be8f7870806071d8fce5a92bc3616333

Observation b797cd48-dee6-449b-8469-a669a19acbf1 · inbound

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models cites this paper.

Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:48:56.083452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:48:56.083452Z digest=sha256:9552ab988ce55a9cab68cddd01747a0122a088bbd072d32e8a983c0ac5fa7c21