Pith. sign in

Paper Citation Record · LEDGER

What If We Recaption Billions of Web Images with LLaMA-3?

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.08478.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08478 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:26:35.252979Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.777451Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8f9c9d56-a50f-40e0-886d-c723d7ea8aed · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.869347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:23ff80c3cc08e1e54b62ef24e926d9d2820d84cf422a505d5d4c6612c141f55f

Observation 4337a5b9-0985-43a8-89b6-cff4964300a7 · inbound

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning cites this paper.

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning What If We Recaption Billions of Web Images with LLaMA-3?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:35.252979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:35.252979Z digest=sha256:24ab7bb8a23a180939cdbc91f336f17b8e746eb403e151717f77a4662fdc7cd3

Observation bca85c64-9114-4285-a012-c084a8a858ba · inbound

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models cites this paper.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.400816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.400816Z digest=sha256:eb1a1240273aa8a43ea54b5c147055f1ef7efd5af8669adbbdde4886b55f3725

Observation 6c41d5a6-4e76-40c6-adc6-d2060604eea6 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors What If We Recaption Billions of Web Images with LLaMA-3?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.478225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.478225Z digest=sha256:d8780698890a2b42baa7a720e2109f87403899efd41c5856e80b8768dd7bbc79

Observation 2124472e-1761-4fa3-88f0-08854232d528 · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.506534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.506534Z digest=sha256:8baebb3d0836972070edc31f417983c7f99d45415691cfefe503269573356482

Observation 7d6bf0c8-dd73-4530-9f8d-6fdb9dc9c34e · inbound

EmoCtrl: Controllable Emotional Image Content Generation cites this paper.

EmoCtrl: Controllable Emotional Image Content Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:38:20.764959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T19:35:52.111730Z digest=sha256:cbf28d5e54dc1333cd310af3c34a9d43eb304e5b1d2ee1d07dd6a892b89c50ce

Observation 00cb32ef-8ef2-4f87-82cc-ab0da037548a · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP What If We Recaption Billions of Web Images with LLaMA-3?

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.779014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:c2ced59820b73cf3a3ff6ab5ea8e563a6c37757517f51ebd8bc052c18bec0bdd

Observation f12471f3-3738-490c-a776-f092d4b9993c · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.745660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:6fb5c2ebe9ad45dad5682eaf09ea47f64c031e4f12d6d27590767149ba75be61

Observation 4f5c73a0-66bb-41d7-b1b2-e33e7ea1a405 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.986368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:41be2c3b10bbf88c89eca05fe1ea272c4fee43e1386bb00d9a45028fb06a037b