Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.08478.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:26:35.252979Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:39:50.777451Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8f9c9d56-a50f-40e0-886d-c723d7ea8aed · inbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4337a5b9-0985-43a8-89b6-cff4964300a7 · inbound
LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning What If We Recaption Billions of Web Images with LLaMA-3?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca85c64-9114-4285-a012-c084a8a858ba · inbound
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c41d5a6-4e76-40c6-adc6-d2060604eea6 · inbound
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors What If We Recaption Billions of Web Images with LLaMA-3?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2124472e-1761-4fa3-88f0-08854232d528 · inbound
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6bf0c8-dd73-4530-9f8d-6fdb9dc9c34e · inbound
EmoCtrl: Controllable Emotional Image Content Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00cb32ef-8ef2-4f87-82cc-ab0da037548a · inbound
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP What If We Recaption Billions of Web Images with LLaMA-3?
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f12471f3-3738-490c-a776-f092d4b9993c · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f5c73a0-66bb-41d7-b1b2-e33e7ea1a405 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.