Pith. sign in

Paper Citation Record · LEDGER

CapsFusion: Rethinking Image-Text Data at Scale

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2310.20550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.20550 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:08:08.450760Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.507046Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b687545a-f19e-447e-9c3a-50662f4f4cad · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding CapsFusion: Rethinking Image-Text Data at Scale

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.751043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:c45a0ace8479581c8d1b4f295273bf0b3ec022423800bfb1540402cf1b253df7

Observation 2b097277-fe44-431b-a487-4ff794e18289 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation CapsFusion: Rethinking Image-Text Data at Scale

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.115040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:937a5e6fba5becdc40ed8e26a42c402762eae323ccbcea4ddd0dc09fc09b94f9

Observation a7276b04-4dc2-40a0-9297-790122fa7487 · inbound

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding cites this paper.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding CapsFusion: Rethinking Image-Text Data at Scale

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.488355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:c371a41d83601cc70cbd01c2a0ccdfb080d07493bdac2607704de61ee61f5900

Observation 282cb1db-1d9d-4e8f-86a7-92e62621b0dd · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CapsFusion: Rethinking Image-Text Data at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.450760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.450760Z digest=sha256:f4053b5d8765c3514ceda6875bbbbbb56bf60ac87bbfbe33cbe9198d33f7608d

Observation 85b5c121-2c3c-4e77-9e74-d93c9127a5d4 · inbound

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models cites this paper.

HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models CapsFusion: Rethinking Image-Text Data at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T11:46:45.543788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:46:45.543788Z digest=sha256:9f110ea1acc25ae6cedca0ede0a9c42efd1616946b8d53c22701092109690f46

Observation 3191fa67-74b5-4f33-b0bd-8457f0ec7b40 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data CapsFusion: Rethinking Image-Text Data at Scale

Reference 259

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:36.678781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:36.678781Z digest=sha256:bcdf6a615c5ea07385216a37aed4ea4de785acd7b5cce03272f7af0dd9680558

Observation b3543965-c233-40b3-83c7-965407277f72 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning CapsFusion: Rethinking Image-Text Data at Scale

Reference 250

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.508435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:86b2b81e3e1eea428d5a810cc4cb09b6e46e235ea32400bb312771944d18eed2