Pith. sign in

Paper Citation Record · LEDGER

Multiscale Vision Transformers

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2104.11227.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.11227 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:43:30.849424Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

56
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f0d87a41-ce01-46e2-bcab-4827e813cb0d · inbound

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers cites this paper.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.849424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.849424Z digest=sha256:049680937e9adbde3e4a97456aca68c9740718ff0aeb4b534728a76b9f7ee8b6

Observation 3f4d0af7-febb-42fc-8b62-73b8ddfc7ac5 · inbound

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition cites this paper.

Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:42.379339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:42.379339Z digest=sha256:314dbc4a270371e2248dd60a55997325373669760bf18374e1b61808a451ffaf

Observation 8526b7d1-b4fe-40b6-944d-7d1e3657a067 · inbound

DEFEND: A Large-scale 1M Dataset and Foundation Model for Tobacco Addiction Prevention cites this paper.

DEFEND: A Large-scale 1M Dataset and Foundation Model for Tobacco Addiction Prevention Multiscale Vision Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T18:34:02.990357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:34:02.990357Z digest=sha256:4a1ffec3ff136032f930723abe9ad53985a5ee27da486b0d71d0caa63c5004da

Observation e89d6ad3-b568-453e-866c-041368fe98c7 · inbound

TabICL: A Tabular Foundation Model for In-Context Learning on Large Data cites this paper.

TabICL: A Tabular Foundation Model for In-Context Learning on Large Data Multiscale Vision Transformers

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:35:02.145880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T13:35:02.018244Z digest=sha256:2261624330508130f17ec0d2394bd814f981ec5873801065d9c544df7dfbdbea

Observation 22a43343-5c80-46be-b353-e8b2756ea021 · inbound

Time-Scaling State-Space Models for Dense Video Captioning cites this paper.

Time-Scaling State-Space Models for Dense Video Captioning Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.102006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.102006Z digest=sha256:10cdb179062b3c41d7ea1e5015e6e4412fa50e9d08e56a43176d90a640e5e610

Observation 704b4fe3-874c-4e3b-8cd9-9fb4c2f7e189 · inbound

Binge Watch: Reproducible Multimodal Benchmarks Datasets for Large-Scale Movie Recommendation on MovieLens-10M and 20M cites this paper.

Binge Watch: Reproducible Multimodal Benchmarks Datasets for Large-Scale Movie Recommendation on MovieLens-10M and 20M Multiscale Vision Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:51:40.930317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:51:40.930317Z digest=sha256:d2c3085757460938d3d6fc737f2e9d8570d8a9f4223039c0de2684d6f0bc3c48

Observation abec9720-0c1a-4b1a-b665-a4e175880518 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding Multiscale Vision Transformers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.413541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:ab4f0174704e097f1b92d3836dd90d1992ba13b1172d540edc8b3cf53f676282

Observation d29c174f-9cf5-4bee-af2c-304a4f449c31 · inbound

LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions cites this paper.

LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions Multiscale Vision Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.406828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T09:29:07.142013Z digest=sha256:b4a948cafc1ce9b48d924953f79476d223dacb41e6000dd7c5024f71a2c9585d