Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2406.08478.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:40:27.679528Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:39:50.777451Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 40b79ca7-7d68-4d30-ae59-f2fc638f8802 · inbound
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions What If We Recaption Billions of Web Images with LLaMA-3?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5558870-ea36-4d87-b501-46ec0d11110b · inbound
Active Data Curation Effectively Distills Large-Scale Multimodal Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d3d568-e7c7-4f6f-be2f-f0b11e9642bb · inbound
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa538cf-a18f-4f5a-8df9-c3f91b7a724b · inbound
Causal Graphical Models for Vision-Language Compositional Understanding What If We Recaption Billions of Web Images with LLaMA-3?
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a0ea1f-5740-4cc8-94f1-31d0b040d747 · inbound
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities What If We Recaption Billions of Web Images with LLaMA-3?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aacd713a-7890-45b6-9c77-726fbabf4219 · inbound
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19cb48d-4b04-4567-aa31-da0a91370c34 · inbound
Dual Diffusion for Unified Image Generation and Understanding What If We Recaption Billions of Web Images with LLaMA-3?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ccd5bd-2c60-4193-b74e-29a7c531c02b · inbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness What If We Recaption Billions of Web Images with LLaMA-3?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1262d9-2858-43e1-ad4d-6566f6f7459e · inbound
Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cfe6054-058c-4f67-ae63-ad829233e3ad · inbound
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e191db48-ad61-4223-9dab-3e802c457a8f · inbound
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ffc7346-fa6e-43a9-a657-849eb4326dfb · inbound
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b97f22f-fc17-494e-b3a8-b910bedb0723 · inbound
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP What If We Recaption Billions of Web Images with LLaMA-3?
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e8f1002-51fa-43a0-9ba5-ee428364b6ee · inbound
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea621ca-859b-4dfe-9e18-a166c150d069 · inbound
Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale What If We Recaption Billions of Web Images with LLaMA-3?
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f9c9d56-a50f-40e0-886d-c723d7ea8aed · inbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 32e8331b-4d92-4f87-b964-32d1cc340a8d · inbound
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations What If We Recaption Billions of Web Images with LLaMA-3?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e30366c-0674-43ef-9b1e-88fd609466fe · inbound
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) What If We Recaption Billions of Web Images with LLaMA-3?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653cfb4f-4adf-40cd-a660-cf1a1f3e2e49 · inbound
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a10dd858-9346-4875-933a-7566f1fd5aac · inbound
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4337a5b9-0985-43a8-89b6-cff4964300a7 · inbound
LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning What If We Recaption Billions of Web Images with LLaMA-3?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 473a62d3-b1fd-4d9b-b6e9-383c663df6d5 · inbound
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality What If We Recaption Billions of Web Images with LLaMA-3?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca85c64-9114-4285-a012-c084a8a858ba · inbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c41d5a6-4e76-40c6-adc6-d2060604eea6 · inbound
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors What If We Recaption Billions of Web Images with LLaMA-3?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2124472e-1761-4fa3-88f0-08854232d528 · inbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6bf0c8-dd73-4530-9f8d-6fdb9dc9c34e · inbound
EmoCtrl: Controllable Emotional Image Content Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 00cb32ef-8ef2-4f87-82cc-ab0da037548a · inbound
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP What If We Recaption Billions of Web Images with LLaMA-3?
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f12471f3-3738-490c-a776-f092d4b9993c · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f5c73a0-66bb-41d7-b1b2-e33e7ea1a405 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0e1bc262-f6f4-4ff5-af74-40febf6090a1 · inbound
Towards Physics-Faithful Generation of Scientific Diagrams What If We Recaption Billions of Web Images with LLaMA-3?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.