Pith. sign in

Paper Citation Record · LEDGER

What If We Recaption Billions of Web Images with LLaMA-3?

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2406.08478.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08478 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:40:27.679528Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.777451Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 40b79ca7-7d68-4d30-ae59-f2fc638f8802 · inbound

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions cites this paper.

CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions What If We Recaption Billions of Web Images with LLaMA-3?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:57:11.884915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:57:11.884915Z digest=sha256:882a80d01bab0404ba459b7c01d701707c705b6dc6b6afb7bd661947543a68a5

Observation f5558870-ea36-4d87-b501-46ec0d11110b · inbound

Active Data Curation Effectively Distills Large-Scale Multimodal Models cites this paper.

Active Data Curation Effectively Distills Large-Scale Multimodal Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T11:06:36.023274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:06:36.023274Z digest=sha256:970368f5af3b9e143afada656a10048b051f0a7eeb0b5db95ab6421633cceb29

Observation a2d3d568-e7c7-4f6f-be2f-f0b11e9642bb · inbound

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models cites this paper.

AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:25:25.549423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:25:25.549423Z digest=sha256:a395026885d4ce85d5fcffd71da397131a0f9e67a8ef78a3e0842940270c4347

Observation 4fa538cf-a18f-4f5a-8df9-c3f91b7a724b · inbound

Causal Graphical Models for Vision-Language Compositional Understanding cites this paper.

Causal Graphical Models for Vision-Language Compositional Understanding What If We Recaption Billions of Web Images with LLaMA-3?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:44.342559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:10:44.342559Z digest=sha256:605a19b4565e6aece59d8ce3f1328efa2b5fc3db692881b3348dd3c1e86c3c61

Observation d2a0ea1f-5740-4cc8-94f1-31d0b040d747 · inbound

UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities cites this paper.

UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities What If We Recaption Billions of Web Images with LLaMA-3?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:59:08.943367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:59:08.943367Z digest=sha256:d154a02ca07c9281fdc32b6ee056b07d723e486d7dd060dcec60748d6ac8d4c6

Observation aacd713a-7890-45b6-9c77-726fbabf4219 · inbound

Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation cites this paper.

Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:14:31.868131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:14:31.868131Z digest=sha256:dbc7afadcbafeeb1af533a5745e101720e551df5510fb558cdc0d44e116a635b

Observation c19cb48d-4b04-4567-aa31-da0a91370c34 · inbound

Dual Diffusion for Unified Image Generation and Understanding cites this paper.

Dual Diffusion for Unified Image Generation and Understanding What If We Recaption Billions of Web Images with LLaMA-3?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.743838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.743838Z digest=sha256:50a57761f96538c5cb7c5ab4c5a8992192a5dee0b789ace5738cf4be3328ca0c

Observation b7ccd5bd-2c60-4193-b74e-29a7c531c02b · inbound

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness cites this paper.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness What If We Recaption Billions of Web Images with LLaMA-3?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.564806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.564806Z digest=sha256:a676844613a0e4dc9e5fde660dc8d1fc50788ecc89af82a587540e110654ec7a

Observation 7d1262d9-2858-43e1-ad4d-6566f6f7459e · inbound

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation cites this paper.

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T11:52:14.792125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:52:14.792125Z digest=sha256:322980d787b8fe8d6b56de54d0219e09f675a2be1adcdad667f3d43cb5d30a19

Observation 0cfe6054-058c-4f67-ae63-ad829233e3ad · inbound

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation cites this paper.

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:41:44.450740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:41:44.450740Z digest=sha256:66fab45625422987b5b2cca60c8cee29299cb2916335c27167f6b79881b57b9a

Observation e191db48-ad61-4223-9dab-3e802c457a8f · inbound

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models cites this paper.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.679528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.679528Z digest=sha256:36d78c6495d0a6c6f4f109807acbf26d831f071f6f714dfe11b5977e0363c8d1

Observation 8ffc7346-fa6e-43a9-a657-849eb4326dfb · inbound

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning cites this paper.

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:29:57.320658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:29:57.320658Z digest=sha256:9e8b2637794f3f82ed969351353fdf9804861af364484bd3326cec876749007f

Observation 8b97f22f-fc17-494e-b3a8-b910bedb0723 · inbound

X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP cites this paper.

X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP What If We Recaption Billions of Web Images with LLaMA-3?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:31.945597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:31.945597Z digest=sha256:c69d94cf75859bbdd7d377e6da6e866e99787bbaf82b3c1a5d01c5bb7d8778cd

Observation 1e8f1002-51fa-43a0-9ba5-ee428364b6ee · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:13.257703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:13.257703Z digest=sha256:f829773fda7f171856bfc9f31c9aca53243897195e2c1c581e63d02ca24d79d9

Observation 3ea621ca-859b-4dfe-9e18-a166c150d069 · inbound

Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale cites this paper.

Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale What If We Recaption Billions of Web Images with LLaMA-3?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T01:05:20.299344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:05:20.299344Z digest=sha256:01277d982fceb526af2c588c3c7f1f5d39051f1087c4539045a6a5393967d7b1

Observation 8f9c9d56-a50f-40e0-886d-c723d7ea8aed · inbound

OmniGen2: Towards Instruction-Aligned Multimodal Generation cites this paper.

OmniGen2: Towards Instruction-Aligned Multimodal Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.869347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:b66b58520fd0ee0e97d5320e2de967dfb61dccf604d3886e10fb383ee41e0298

Observation 32e8331b-4d92-4f87-b964-32d1cc340a8d · inbound

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations cites this paper.

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations What If We Recaption Billions of Web Images with LLaMA-3?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:10.699688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:10.699688Z digest=sha256:a6402c7cee391be367d9b400759eb74d5efce5fd2fd36c5b07f6a1c09d26ac3f

Observation 1e30366c-0674-43ef-9b1e-88fd609466fe · inbound

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) cites this paper.

Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M) What If We Recaption Billions of Web Images with LLaMA-3?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.065213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.065213Z digest=sha256:7325969854f27a8ad4626b0404cc22137e3cf4c12866c707a04e2d555f8e382f

Observation 653cfb4f-4adf-40cd-a660-cf1a1f3e2e49 · inbound

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models cites this paper.

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:54:06.240933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:54:06.240933Z digest=sha256:c94428e3a44a4576021cc4e28d1fc1e3bd9b5db99e3d4079fa2e35ef0f61be00

Observation a10dd858-9346-4875-933a-7566f1fd5aac · inbound

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation cites this paper.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.066023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.066023Z digest=sha256:33f999a61f8ccc187408dc4be8736eebe65f92f85d8e637ac9e4594f34ac8e1b

Observation 4337a5b9-0985-43a8-89b6-cff4964300a7 · inbound

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning cites this paper.

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning What If We Recaption Billions of Web Images with LLaMA-3?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:35.252979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:35.252979Z digest=sha256:d531049a826b36fd16305768d4742cdfc54e2de73cba1a0b47dbaeeacb1e510c

Observation 473a62d3-b1fd-4d9b-b6e9-383c663df6d5 · inbound

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality cites this paper.

Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality What If We Recaption Billions of Web Images with LLaMA-3?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:51:42.843903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:51:42.843903Z digest=sha256:54e40c758c200b4caf6af6a46f33580f24beff13fa5a1abdc99a98e7b476e66a

Observation bca85c64-9114-4285-a012-c084a8a858ba · inbound

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models cites this paper.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.400816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.400816Z digest=sha256:924fff5ad0645e315f8f325acf071995401b3821dffd8d21b2fa51b55ccddcc9

Observation 6c41d5a6-4e76-40c6-adc6-d2060604eea6 · inbound

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors cites this paper.

Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors What If We Recaption Billions of Web Images with LLaMA-3?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:54.478225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:05:54.478225Z digest=sha256:83974b9b20e8b2d2d468c613fecddf22bdb9e2adf9ad08551a7426a2b19db1bc

Observation 2124472e-1761-4fa3-88f0-08854232d528 · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning What If We Recaption Billions of Web Images with LLaMA-3?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:35.506534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:35.506534Z digest=sha256:65790716dc6b374cd51152a2ca79c5b6a70e3badcf640d27e3f97d8f6ed017ff

Observation 7d6bf0c8-dd73-4530-9f8d-6fdb9dc9c34e · inbound

EmoCtrl: Controllable Emotional Image Content Generation cites this paper.

EmoCtrl: Controllable Emotional Image Content Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:38:20.764959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T19:35:52.111730Z digest=sha256:a244e3e2bd1699e4f23e360e2c747f1601c4b7de65c87363367487b205acb8e4

Observation 00cb32ef-8ef2-4f87-82cc-ab0da037548a · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP What If We Recaption Billions of Web Images with LLaMA-3?

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.779014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:52d1c8118d68467f7b8d401557975425335f96e8dc8a2b0a36b1be46a6af267e

Observation f12471f3-3738-490c-a776-f092d4b9993c · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.745660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:0fcff4eea938e67431a3ba40458d268c78a0bb88848e6d3074b143314377eeb8

Observation 4f5c73a0-66bb-41d7-b1b2-e33e7ea1a405 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models What If We Recaption Billions of Web Images with LLaMA-3?

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.986368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e81535010a7f7f29246fc1d8459c74865dcdab8eb8ee2f3f3e8e5a40a619a216

Observation 0e1bc262-f6f4-4ff5-af74-40febf6090a1 · inbound

Towards Physics-Faithful Generation of Scientific Diagrams cites this paper.

Towards Physics-Faithful Generation of Scientific Diagrams What If We Recaption Billions of Web Images with LLaMA-3?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:36:33.735190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:36:33.735190Z digest=sha256:1408b96b8d0ec571f18cebb928aeb4182b22ec73b1c0e55ac3e80868e85c3ac5