Pith. sign in

Paper Citation Record · LEDGER

WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2103.01913.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2103.01913 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:37:15.327747Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:50.767139Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 19d97322-a6a2-4e77-a914-2cd1e5d21699 · inbound

Sigmoid Loss for Language Image Pre-Training cites this paper.

Sigmoid Loss for Language Image Pre-Training WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:05:36.559014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T13:05:36.460932Z digest=sha256:0415350d0016bbc7abc64a8dabdec64c208ecff4c88e5952556c32bc24f2372d

Observation a4349020-5751-42d9-aa17-5b7c0ea28b12 · inbound

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images cites this paper.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.327747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.327747Z digest=sha256:32461f8ca88b9c33c8f5344ae4b83094add3c91ced8f866d2d3041495ac71c41

Observation d30ae5c1-d8f0-4773-b744-a8f70748e87a · inbound

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning cites this paper.

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:34:40.435496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:34:40.435496Z digest=sha256:74f4d3a174713635d30067c2128344002d936491f56adc9b4e2ee47169dc87ae

Observation 0f035a0c-dd9e-4a42-9e26-7c81387272fc · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:50.768586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:a862b8bff9188440bf73c1b66037fc19e76b8ca80117b22ba0bc744f4458e8a2