Pith. sign in

Paper Citation Record · LEDGER

CoAtNet: Marrying Convolution and Attention for All Data Sizes

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2106.04803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.04803 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:19:33.862479Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T11:41:02.799655Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1e08dccc-3a69-42d3-b1e7-7942c5e5b7ba · inbound

MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer cites this paper.

MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:46:35.211861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T20:46:35.073600Z digest=sha256:fec3bed5a1a60976b60a7c18e6690b1fe20f74602dc073830dfbc4bfe24c6678

Observation bd6b2b45-d640-41e6-972f-b6f4603c7867 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.485888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:509aa3f2c845fdf9632540896a0367cf1e01589158fb4aef4bb6c18539f298d0

Observation 07410519-75ec-4f48-99b6-0727dd33450b · inbound

DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection cites this paper.

DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T11:19:33.862479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:19:33.862479Z digest=sha256:7b1d3446a13b0a1bef3c5a0c8bd5ed79754c6a1ac62d3540d545dc589db3a1e4

Observation 2c9c4fae-ac83-4d4e-8b8e-4d89912c42c4 · inbound

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification cites this paper.

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:45.974828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:21:45.974828Z digest=sha256:e0555f25949d71c642462a407d8e126ee57f3f312167a2aae68145d85f785f39

Observation 3c15361b-e1a1-4a4e-bef6-28156de9a622 · inbound

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents cites this paper.

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.801223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-05T11:32:36.356530Z digest=sha256:02055caf031ab4291b8afdd828bb2c540463e9b775a2c110f11d3465f29b63f1

Observation 30f27239-2882-49e5-bec0-8f7fd7a4ed2c · inbound

Advancing Vision Transformer with Enhanced Spatial Priors cites this paper.

Advancing Vision Transformer with Enhanced Spatial Priors CoAtNet: Marrying Convolution and Attention for All Data Sizes

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.664602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T05:22:21.264807Z digest=sha256:157f4004462ffec8cfc2efca1b063d2a0e6a23ae29155565e1ac6b3435dcc806