Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T09:38:09.427509Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 41 inbound Pith citation observations for arXiv:2111.11432.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T09:38:09.427509Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T22:15:39.116775Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T20:00:08.247992Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 646eacc5-fe24-4a40-8e45-d12ed178fc53 · outbound
Florence: A New Foundation Model for Computer Vision W., Alexander, M
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ee358cf-341d-4271-81d1-4e8c34f83c33 · outbound
Florence: A New Foundation Model for Computer Vision On the Opportunities and Risks of Foundation Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e7085aab-de62-453a-943b-9d4594c94276 · outbound
Florence: A New Foundation Model for Computer Vision Language Models are Few-Shot Learners
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98282352-255a-416f-9d0a-266912a68955 · outbound
Florence: A New Foundation Model for Computer Vision Learning the Best Pooling Strategy for Visual Semantic Embedding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd6b2b45-d640-41e6-972f-b6f4603c7867 · outbound
Florence: A New Foundation Model for Computer Vision CoAtNet: Marrying Convolution and Attention for All Data Sizes
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0742c473-51f0-4d0d-8fe9-b1fb1ae8318e · outbound
Florence: A New Foundation Model for Computer Vision BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0159993e-eb37-427d-b09a-5c9a47231907 · outbound
Florence: A New Foundation Model for Computer Vision CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 744b6ce5-f4c4-44e9-afb1-9ddf3d261a1c · outbound
Florence: A New Foundation Model for Computer Vision An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 883c8e61-3f00-46d1-972a-caa11b411eb8 · outbound
Florence: A New Foundation Model for Computer Vision Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 12135bce-d502-4a78-a3c0-60d960fac5d8 · outbound
Florence: A New Foundation Model for Computer Vision Rich fea- ture hierarchies for accurate object detection and semantic segmentation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6085eebd-ed68-41ce-b8dc-6c865a7e665b · outbound
Florence: A New Foundation Model for Computer Vision Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8378e823-52a4-41e8-ae91-90238dadc7c4 · outbound
Florence: A New Foundation Model for Computer Vision Big Transfer (BiT): General Visual Representation Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94108b8c-708d-40b2-bddb-67f32fdfbe15 · outbound
Florence: A New Foundation Model for Computer Vision Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f9c80622-dd4a-4296-8644-c964521a19a2 · outbound
Florence: A New Foundation Model for Computer Vision Video Swin Transformer
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 634c8446-581f-4fc9-a4e7-62490b733990 · outbound
Florence: A New Foundation Model for Computer Vision Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a244ce7a-635c-43e8-a289-496097e726dc · outbound
Florence: A New Foundation Model for Computer Vision ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc8e0c78-d932-4e80-8598-c0117f88e847 · outbound
Florence: A New Foundation Model for Computer Vision Learning Transferable Visual Models From Natural Language Supervision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8fccb61-5695-4904-b2a8-14516ad85837 · outbound
Florence: A New Foundation Model for Computer Vision Zero-Shot Text-to-Image Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5d6fb1a-b50a-43e6-b7e3-d0504daab814 · outbound
Florence: A New Foundation Model for Computer Vision TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1f6e3149-64d8-4ff5-a89f-371d32ac1f14 · outbound
Florence: A New Foundation Model for Computer Vision MiniVLM: A Smaller and Faster Vision-Language Model
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c3d75b6a-ad9f-4644-8aa6-e8b6db3a6a98 · outbound
Florence: A New Foundation Model for Computer Vision ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c79ce03d-e652-4e1a-9612-2a20441e7ebf · outbound
Florence: A New Foundation Model for Computer Vision SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c368a540-f6e6-4f13-b914-96d9a1d803bd · outbound
Florence: A New Foundation Model for Computer Vision Focal Self-attention for Local-Global Interactions in Vision Transformers
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7589c8ec-0140-4b7f-81b2-ff07c748df29 · outbound
Florence: A New Foundation Model for Computer Vision FILIP: Fine-grained Interactive Language-Image Pre-Training
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a535edb2-3427-45e4-b3af-0912d9ddb31d · outbound
Florence: A New Foundation Model for Computer Vision ERNIE-ViL: Knowledge Enhanced Vision-Language Representations Through Scene Graph
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0b35262-610b-473b-b275-7e8eaa5a55e4 · outbound
Florence: A New Foundation Model for Computer Vision Scaling Vision Transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b409af3-6a61-4058-8f32-c7eae8f533be · outbound
Florence: A New Foundation Model for Computer Vision Simple multi-dataset detection
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c16ad677-20fd-41bd-936e-294af517c9d2 · outbound
Florence: A New Foundation Model for Computer Vision D., and Le, Q
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2606703f-208d-4c54-b173-15f0bc8717ff · inbound
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection Florence: A New Foundation Model for Computer Vision
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6cb1ad25-4b62-4acd-b798-5e84af0ca8e4 · inbound
Flamingo: a Visual Language Model for Few-Shot Learning Florence: A New Foundation Model for Computer Vision
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b40b5527-813c-4b27-a16f-ddeaf2dd9cbd · inbound
CoCa: Contrastive Captioners are Image-Text Foundation Models Florence: A New Foundation Model for Computer Vision
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 475ecc57-76ba-4766-9cee-d2ca7480e9af · inbound
GIT: A Generative Image-to-text Transformer for Vision and Language Florence: A New Foundation Model for Computer Vision
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d23bfd48-e92b-49ec-9cb4-7a41a944438e · inbound
DetailCLIP: Injecting Image Details into CLIP's Feature Space Florence: A New Foundation Model for Computer Vision
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 68086a93-8ac9-4b82-af55-9ed2e72b5cd9 · inbound
PaLI: A Jointly-Scaled Multilingual Language-Image Model Florence: A New Foundation Model for Computer Vision
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 109f2a8e-3b42-4a52-a230-15d8417d5650 · inbound
InternVideo: General Video Foundation Models via Generative and Discriminative Learning Florence: A New Foundation Model for Computer Vision
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 062bf9a8-dc32-4144-b7b8-d670d279f8ae · inbound
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models Florence: A New Foundation Model for Computer Vision
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cf475cc-c32a-409d-bf65-ba2d85ad209f · inbound
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action Florence: A New Foundation Model for Computer Vision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aba21aa4-6968-4a88-96a4-b0b78cbc870f · inbound
Sigmoid Loss for Language Image Pre-Training Florence: A New Foundation Model for Computer Vision
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f31541d-d06c-400f-bf8b-1c820c1cfb8e · inbound
Visual Instruction Tuning Florence: A New Foundation Model for Computer Vision
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 33860558-ca6c-4577-a7c3-470807f3a221 · inbound
VideoChat: Chat-Centric Video Understanding Florence: A New Foundation Model for Computer Vision
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 57929d05-82c6-4922-bb9c-552f9b78c597 · inbound
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models Florence: A New Foundation Model for Computer Vision
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a20c6a98-1483-4d79-9f9a-7c562805434f · inbound
Demystifying CLIP Data Florence: A New Foundation Model for Computer Vision
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 370c499d-c546-4fe6-a5be-2eeb257b052f · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Florence: A New Foundation Model for Computer Vision
Reference 172
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f589d9b-43fd-44f4-b4d8-c27517ce5713 · inbound
LLaVA-Video: Video Instruction Tuning With Synthetic Data Florence: A New Foundation Model for Computer Vision
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51d341cc-7c45-4066-bc65-32020cc8e307 · inbound
Interactive Program Synthesis for Modeling Collaborative Physical Activities from Narrated Demonstrations Florence: A New Foundation Model for Computer Vision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 465ef22e-5fa9-40d3-912c-24d9960e4c24 · inbound
BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning Florence: A New Foundation Model for Computer Vision
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdfc8ab6-21f0-4117-b4f9-373cb95d0849 · inbound
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Florence: A New Foundation Model for Computer Vision
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ba6292-fb13-46e6-ada7-9f657d7f01cc · inbound
GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models Florence: A New Foundation Model for Computer Vision
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6b04895-2c6a-4c50-93b5-db3ac162f866 · inbound
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining Florence: A New Foundation Model for Computer Vision
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0dc396d5-ad68-4320-860c-2d05b2ccebba · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Florence: A New Foundation Model for Computer Vision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 78ed5de2-c5bb-408d-b333-1c9d14c3adc7 · inbound
WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Florence: A New Foundation Model for Computer Vision
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 735e340d-7e27-4998-806f-c5c5bfd6f444 · inbound
Omni-NegCLIP: Enhancing CLIP with Front-Layer Contrastive Fine-Tuning for Comprehensive Negation Understanding Florence: A New Foundation Model for Computer Vision
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 494ca35b-4f19-495e-9ede-23a4ea699821 · inbound
Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook Florence: A New Foundation Model for Computer Vision
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 245f6a2a-1b77-49ec-a9b1-12293709a5f8 · inbound
From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media Florence: A New Foundation Model for Computer Vision
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d481afe7-dc6b-4d57-b5aa-7bb4d851c42a · inbound
Split and Aggregation Learning for Foundation Models Over Mobile Embodied AI Network (MEAN): A Comprehensive Survey Florence: A New Foundation Model for Computer Vision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 78cb5385-fa9d-4ff1-8ab6-5f8dae6f858e · inbound
Joint Semantic Token Selection and Prompt Optimization for Interpretable Prompt Learning Florence: A New Foundation Model for Computer Vision
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f055904-4c87-41ea-9649-f978570c8775 · inbound
Cluster-Aware Neural Collapse Prompt Tuning for Long-Tailed Generalization of Vision-Language Models Florence: A New Foundation Model for Computer Vision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01b8d51e-c4b3-4d69-879b-73002cb13525 · inbound
Text-Guided Visual Representation Learning for Robust Multimodal E-Commerce Recommendation Florence: A New Foundation Model for Computer Vision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6bc691eb-3df3-4163-8746-341d33544a0c · inbound
Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics Florence: A New Foundation Model for Computer Vision
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 44aecb47-efe3-43b2-a692-70290b73dddd · inbound
The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP Florence: A New Foundation Model for Computer Vision
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 22079198-aca6-4773-9ec7-dcefec27487f · inbound
CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning Florence: A New Foundation Model for Computer Vision
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ea42d82-df68-432a-891f-ff8252dc7af2 · inbound
Count Anything Florence: A New Foundation Model for Computer Vision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7cfe39c1-2707-4331-bd1f-3d7f280fa569 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Florence: A New Foundation Model for Computer Vision
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1cc72dca-4e34-4c94-afef-9e6e13e073ee · inbound
TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition Florence: A New Foundation Model for Computer Vision
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95064b9b-dde1-49b7-a9fb-5cb1b877550f · inbound
TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition Florence: A New Foundation Model for Computer Vision
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 783b0d80-4a18-44a1-aec5-1681b6494703 · inbound
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception Florence: A New Foundation Model for Computer Vision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db37b18-b123-4b86-9e1c-c85f333921a7 · inbound
Qwen-Audio-VAE Technical Report Florence: A New Foundation Model for Computer Vision
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a7e214-2868-47c6-b13b-c8886085bde8 · inbound
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Florence: A New Foundation Model for Computer Vision
Reference 219
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee77af0-edae-4a26-98d3-02c0a9cae49b · inbound
Lexical discovery in unknown environments orchestrated by Large Language Models Florence: A New Foundation Model for Computer Vision
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.