Pith. sign in

Paper Citation Record · LEDGER

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2001.07966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2001.07966 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:38.760337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:38:09.540556Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a244ce7a-635c-43e8-a289-496097e726dc · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.543090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:d37c8d301c5ae73e4279c401ce2755908c971495498a84ee21b5d903c3bb61e8

Observation ca36181f-c706-41ec-8116-a6c34923517b · inbound

CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance cites this paper.

CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:00.109020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:00.109020Z digest=sha256:e24069a119af9921a3588a52f900683d0547aa02b745073671602ecbdcb44775

Observation e6158d28-b18e-4d1d-a4b9-87edc4598691 · inbound

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples cites this paper.

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T16:32:24.428017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:32:24.428017Z digest=sha256:05b28d8a5961cdb4ed8e94d6b1158eb242c0d933c6037c2c79890be46fc038b3

Observation 83ece59c-1713-4f08-9db8-d05488a96e43 · inbound

GeoMM: On Geodesic Perspective for Multi-modal Learning cites this paper.

GeoMM: On Geodesic Perspective for Multi-modal Learning ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.760337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.760337Z digest=sha256:b5bd03609d7bd15f5a3be37bd71d4d28d398b8c0b0b484314c9ad02144b533fc