Pith. sign in

Paper Citation Record · LEDGER

MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2112.01526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.01526 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:18:52.587351Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.067076Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 994f06a4-e599-46e3-a9c0-1ec28f86da27 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:46:09.985333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:c0de0d567c1044a2f20ede712116df2d0d8cb4cd5636a59134f75b8cdde1bb77

Observation 39f2950a-f402-4421-b909-5b173f8fea03 · inbound

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery cites this paper.

STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:52.497930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:52.497930Z digest=sha256:5a63ae47bd2d15eb86b19bc18b240a56d40f74c6db515032a24eafb993f3a956

Observation 47667847-793a-41e5-a300-01a84560394c · inbound

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer cites this paper.

Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:01.549104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:07:41.426142Z digest=sha256:114bb2c48bf3eea436baf0ce74b1847cf849bf487780c370b282923569c34da2

Observation d0144242-2be9-4ce4-be83-9dee0d1a0f39 · inbound

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition cites this paper.

SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:30.947512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:23:07.258575Z digest=sha256:2477b646751cce728ddcee48a36e464c057b90385b5334c479d5b4612505ffee

Observation e0894763-fb72-4958-9a10-3107a2386b37 · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.068937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:67835e794315fc9ecdd9fb31987fda870bab12ac18eba6c0803645bf8139a9b9

Observation 6dfce47d-8805-492e-b407-0cba9d26e496 · inbound

Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective cites this paper.

Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T23:58:47.097757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T23:58:47.097757Z digest=sha256:e9c78ea9c5b084594a9677f15370bdf5d7c43f7d28fe9273a49fce5dec4a46f8

Observation 3800e9ee-fe5c-4b74-bba0-876295b59135 · inbound

Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines cites this paper.

Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines MViTv2: Improved Multiscale Vision Transformers for Classification and Detection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:52.587351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:18:52.587351Z digest=sha256:6e0bfcd55b9e9a6653e6be0b88543e233453cb262e6ae384a443a2ee728c6a70