Pith. sign in

Paper Citation Record · LEDGER

Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2004.00849.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2004.00849 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T20:29:13.300410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:17:07.110860Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 546fb144-4606-4b62-9fba-c2f0036bb45a · inbound

GIT: A Generative Image-to-text Transformer for Vision and Language cites this paper.

GIT: A Generative Image-to-text Transformer for Vision and Language Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:54:07.656061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T20:54:07.572136Z digest=sha256:6de992704579d97a9e825290a8ea2a0201933d6b78ec6e85965e722ce977c1e0

Observation f8337825-ee2a-41f7-b9bf-622fe1072f80 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:26:06.416304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:7b41e5e219bfd9ed280a13a00bb642cc17f529cdfc2ffae78f2959e03fdc4067

Observation 41625290-b7e6-4102-898a-aaf729961f51 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:27:59.078329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:3515ece6be0e23c7661179346397a23687a45fc0ed2a9f81baa84e39e90aca1d

Observation 17b0b189-121f-48ae-9ba4-4a09a08e92c7 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 287

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:25:59.812425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:a666d1365e53f829fc649dc0bef8ed387fa0211403c331f5c91f4e96b1fe5e8e

Observation b271ee9a-b7ac-4699-8be2-4a2b025c1f69 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:39:22.424251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:814d03e22ca4650b0e6022edf183345286ea76c2a794805d6bfa8358692791f6

Observation e62cc5bd-370a-4985-a8b9-d5f8d555a58b · inbound

Kinky vortons in the 2HDM cites this paper.

Kinky vortons in the 2HDM Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T20:29:13.300410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:29:13.300410Z digest=sha256:687990e549ebe0d0938d7b3c8d2dd9c57c073d496ddceb3e0b4754b92b6a2f6a

Observation 05e19b0a-5d10-4397-bf3b-759b7091ba53 · inbound

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding cites this paper.

HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.113515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-02T15:15:24.784685Z digest=sha256:05b6122528807341615e58eda89adad82a19ce4616cb7557d8cd59d8bf8f3f2d