Pith. sign in

Paper Citation Record · LEDGER

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers

As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2411.18995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18995 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:43:31.987951Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7de7d48-ba92-4b66-8577-9f33223c599b · outbound

This paper cites Layer Normalization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.467766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.467766Z digest=sha256:83ec89ba96bd6728eca51c63a962284e8d7c3f3c9449165a6d42961cd414e2c1

Observation 7475193f-5158-467a-ace5-713f4e0b2708 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.508733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.508733Z digest=sha256:9de9896dc15b34e9aa7c85144473b3dd9363e4530ef0d94d4a4cea8d04d72c14

Observation 10321d41-0599-4d38-80f6-c09ac586fcb8 · outbound

This paper cites CycleMLP: A MLP-like Architecture for Dense Prediction.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers CycleMLP: A MLP-like Architecture for Dense Prediction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.546956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.546956Z digest=sha256:ee82bc89a989233a46af77890d29a8f9d5bb251cf8e31343f0ab5da5a0d47903

Observation 39a28443-e4a1-4509-9530-e1271565b2c6 · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers, 2021.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Twins: Revisiting the design of spatial attention in vision transformers, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:35.070608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.552877Z digest=sha256:259778d9823638cc1f7d93baa51bfb2673f6d2b446f9a7e1f737a939cee571d3

Observation 38fee022-f3ce-46fc-9af3-cb39d6b99425 · outbound

This paper cites Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.627245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.627245Z digest=sha256:b09e88a125ea5b633a94cc577ad1d31343ce60c280156f69c1acef6b856a6c78

Observation d348127d-e7fc-4b8a-84af-5b9fa2414222 · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.917980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.688855Z digest=sha256:7f3765ec873c8057d37192c109e3f3ea95033587814e5cd6beb2c83017dc1f8c

Observation aaa21eac-edf0-464d-87df-3800442065ce · outbound

This paper cites Coatnet: Marrying convolution and attention for all data sizes.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Coatnet: Marrying convolution and attention for all data sizes

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.859205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.738670Z digest=sha256:c69524a5db98d7c57288e3064eece9e3dc9d14797c7ebf0d682d399154203095

Observation 2c4525cc-c989-4802-9513-0e2f03475a57 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Imagenet: A large-scale hierarchical image database

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.832140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.744332Z digest=sha256:519dae58bb57a7d84daedef09a3b447783a84bde40738888928fa44a51128f07

Observation 9841c4be-b33b-4317-beb2-93eade38f88f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.796276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.796276Z digest=sha256:51eb644b7d44eb8e6fc0c6bb245896a1996c8fda2e4bdb78ff855d741976ed14

Observation f0d87a41-ce01-46e2-bcab-4827e813cb0d · outbound

This paper cites Multiscale Vision Transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.849424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.849424Z digest=sha256:a244ee79d4344f36a993faec1fb2a8024d401cc433d74a7ffc6b7a98eecd3397

Observation bd35d36f-b548-49eb-af81-00f068998dad · outbound

This paper cites Segnext: Rethinking convolutional attention design for semantic segmentation,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Segnext: Rethinking convolutional attention design for semantic segmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.726346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.872013Z digest=sha256:9eafd3932ade4a3bfd54f822bda7e9e6a36157c7e721678626c53000d62788f6

Observation e9170c0f-c16c-4480-8fb4-a70225783446 · outbound

This paper cites Visual Attention Network.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Visual Attention Network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.877886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.877886Z digest=sha256:7f6aab0f5abfc3d5c4ecac05c1d2b3e5906b0b12afb1863622d000935321d5f0

Observation b4b86f3a-ccc3-42c6-9343-2aaeb8f003ab · outbound

This paper cites Deep residual learning for image recognition.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Deep residual learning for image recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.617654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.900984Z digest=sha256:fc23ae3729a90c685f345ab4db29d6d5cdaddc95491863e71d44e111715ab907

Observation 1b8d0802-2d77-4adb-93bf-6fad9df567e1 · outbound

This paper cites Mask r-cnn, 2018.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mask r-cnn, 2018

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.602047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.967885Z digest=sha256:3dae130357926b1204801614a4c3890bbcb26de4db5c292d200fcda6dcb846ff

Observation 9a869c6b-0f24-41a6-b7d7-23d3e7b6be31 · outbound

This paper cites Deep networks with stochastic depth.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Deep networks with stochastic depth

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.556396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:30.997621Z digest=sha256:759191b6d4785f6fba27cae085af3ac2117ffe72ac255eab827ee84beb9c4351

Observation bed20438-0074-4926-b6e5-a5a779e30582 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization, 2017.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Arbitrary style transfer in real-time with adaptive instance normalization, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.424854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.038529Z digest=sha256:500e4f27de8d869ac5b920fab0c8b14ce2034d2c83284340dd2e129946dee6a4

Observation 29986528-00a0-4f41-93d0-31d13389fea1 · outbound

This paper cites Batch renormalization: Towards reducing minibatch dependence in batch-normalized models, 2017.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Batch renormalization: Towards reducing minibatch dependence in batch-normalized models, 2017

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.356058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.051916Z digest=sha256:207838f3ff52435ee0f64bcd9b78ae1116754a6bb9f6677a6d074798a5ccad32

Observation e595f8bc-5db7-46cf-a31e-7b9e1280c0fe · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.056554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.056554Z digest=sha256:6e2635ab7098a0f95eacce06b44ba2f97c487768e8d0fb5fc93b27b4d098d39c

Observation 1fa7156f-bf36-4b4d-8030-57372fbb9e40 · outbound

This paper cites Relational self-attention: What’s missing in attention for video understanding.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Relational self-attention: What’s missing in attention for video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.341086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.061886Z digest=sha256:862094487e6319892a2564b22f2eaf25965a7e8bb5eb4a5dce8f78ec95951c2e

Observation f2835c07-1940-4809-b32c-329e8f100596 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.086558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.086558Z digest=sha256:28a59f9c268f261d5543be9e6dc6fdb3f110a574f5c5281bb45c87ef32bc9a05

Observation 5722a313-f18f-455e-935c-0cd05c2e3472 · outbound

This paper cites Panoptic feature pyramid networks.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Panoptic feature pyramid networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.308883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.148497Z digest=sha256:816f5349b12ee5c4fc832994d6cd9924d45b1c81c37d9f4664187303529defc0

Observation a9de85a0-e180-435f-98e7-8c8daa20cef3 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mvitv2: Improved multiscale vision transformers for classification and detection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.154322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.193362Z digest=sha256:064ef878e814971284f793090fad727a7424738743a5ba875f6c3e91feb04a57

Observation d7166ac2-e40f-425b-b744-885ac63c940e · outbound

This paper cites AS-MLP: An Axial Shifted MLP Architecture for Vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers AS-MLP: An Axial Shifted MLP Architecture for Vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.249528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.249528Z digest=sha256:fb540a0659db4ecaefc544f2eb88e499d2d511006c2caadf4b7ce3c4fbfbe490

Observation 3ca758f5-1059-4aae-9d56-02f41a15dc4a · outbound

This paper cites Microsoft coco: Common objects in context.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Microsoft coco: Common objects in context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.273700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.273700Z digest=sha256:67f052c05d8f60784f8dc574c5705bc92f624c83dc3e2e59816c781ac0e6e0c0

Observation 331bfe55-e673-4641-87c1-7b13291d4340 · outbound

This paper cites Focal loss for dense object detection, 2018.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal loss for dense object detection, 2018

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.080426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.278519Z digest=sha256:1ef3c9d3f929072ae006d6c55fafc0cf9a22feac2e6665a1b6db3d9a8022c7ce

Observation 50af3cc3-dbaa-4b87-ae15-d1ae37dfcb64 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.057496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.283982Z digest=sha256:54a85024197211b8595cd1fa0fcd90b91dce4cc6cb4d1c44319c2c85dbfd2336

Observation 5732a23a-d106-486c-afc5-4c97c07cb58a · outbound

This paper cites A convnet for the 2020s.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers A convnet for the 2020s

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.288580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.288580Z digest=sha256:2bd2ba68f57ee8161434a642790624973912dcc252a01facb623c2adb0a87680

Observation 11d4a1a3-db15-4655-913e-64721d343196 · outbound

This paper cites Decoupled Weight Decay Regularization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.345474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.345474Z digest=sha256:665b2b23d32a5f08148e19171b9012854bafe5d1d548f08801928cf69cfbb4d1

Observation 3637f4e0-8dfe-4153-86c9-994b75885459 · outbound

This paper cites How do vision transformers work?, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers How do vision transformers work?, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.920037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.408851Z digest=sha256:448a87ad04428f6473a6a67f691222f0cb90228a43943b1862e49e82905a0da8

Observation 66a09598-7ecb-4a39-b708-bd37f83e4407 · outbound

This paper cites Semantic image synthesis with spatially-adaptive nor- malization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Semantic image synthesis with spatially-adaptive nor- malization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.436305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.436305Z digest=sha256:894ebeeeff0c30274518d91e80496ed7b3d0a5997f3b97edc119c6c122908f47

Observation 1d79b78b-fc4a-46b3-94eb-64f62fc40bc7 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Pytorch: An im- perative style, high-performance deep learning library

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.828294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.484874Z digest=sha256:706e3ecf4205fd0954a86a326d6f05bfed9e12034a7f1fc7c3f521039cb56f1a

Observation f007f04f-8fa1-4923-a4a2-595b4b2d6a9c · outbound

This paper cites Acceleration of stochastic approximation by averaging.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Acceleration of stochastic approximation by averaging

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.813907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.489768Z digest=sha256:42e3130ebd0a6e32bd3995594810f92964facb0d6647f1e5a2b4df14631d61fc

Observation 7ec5e319-f1ad-4b0e-8730-94dc35119152 · outbound

This paper cites What makes for good tokenizers in vision transformer? IEEE Transactions on Pattern Analysis and Machine Intelligence,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers What makes for good tokenizers in vision transformer? IEEE Transactions on Pattern Analysis and Machine Intelligence,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.799813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.494333Z digest=sha256:e19f283459cfc98885b33a4922cee69f9dfbd4f65d0ac1cc6aa5ed955dde85b3

Observation 85eb279b-be1e-4cd5-971b-1d3f423a896c · outbound

This paper cites Do vision trans- formers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34:12116–12128,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Do vision trans- formers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34:12116–12128,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.499973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.499973Z digest=sha256:cbe921c54cb4777a00f21399253dd7a8332bd2a198c829966e1fdfc211409fc7

Observation a066e8f0-2382-4d92-9860-4fdfb750d206 · outbound

This paper cites Mobilenetv2: Inverted residuals and linear bottlenecks, 2019.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mobilenetv2: Inverted residuals and linear bottlenecks, 2019

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.651012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.504794Z digest=sha256:8c4edf22cec968b6a6f4676963a8153e0e1858acdd829400ee98b4674ce05373

Observation 43ce6aa9-f109-4aee-8aa8-3c833791980f · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.561151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.509672Z digest=sha256:5ba315306c6d3bdfcbacbe8cb60c71b359a359f5c65d86fd4d468fb94ae02342

Observation 00b7bd86-6602-4cb8-b4ae-862074991779 · outbound

This paper cites Powernorm: Rethinking batch normaliza- tion in transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Powernorm: Rethinking batch normaliza- tion in transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.546723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.514224Z digest=sha256:7a4aa08ae1880f163336e836b08a737e3dc33bbd00a9ba76e817b3ef3bbc6a1e

Observation 5b13d27e-365c-4453-a5ce-616fe2a4970d · outbound

This paper cites NormFormer: Improved Transformer Pretraining with Extra Normalization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.585543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.585543Z digest=sha256:4d64b6575b0dee616787edd6357a2d8c5bdbea08ab90eb74e7549347f4b37ade

Observation b5ae794c-f344-4b07-9303-94610bc313a6 · outbound

This paper cites Evalnorm: Esti- mating batch normalization statistics for evaluation, 2019.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Evalnorm: Esti- mating batch normalization statistics for evaluation, 2019

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.532254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.661340Z digest=sha256:69607aedf8a27fb07184495b23601dd7c75772e4bdb79349d621277b195624f4

Observation 2dd01271-4d17-4d0b-a9c3-50a8174ddd71 · outbound

This paper cites Rethinking the inception archi- tecture for computer vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Rethinking the inception archi- tecture for computer vision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.689239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.689239Z digest=sha256:c2a4b4da811a464ddd93a0bb2ad891f39fcfca1c5c99644df3fd4d6f61ccbe8d

Observation af4e1c54-1194-4a2e-a0ac-16d053fd602c · outbound

This paper cites Mlp-mixer: An all-mlp architecture for vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mlp-mixer: An all-mlp architecture for vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.693505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.693505Z digest=sha256:6270d91b9c72c200c82fe105662244736975f4379ff98722e778c3fea2bc04cb

Observation 64fd3689-1a74-44e6-9569-c61fbf8cb174 · outbound

This paper cites Resmlp: Feedforward networks for image clas- sification with data-efficient training, 2021.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Resmlp: Feedforward networks for image clas- sification with data-efficient training, 2021

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.465429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.698267Z digest=sha256:8fdecf8b23f5ae5eb823b95db76db4a38791e97815a60a0e9b80f0115b19f501

Observation a2ea657f-b050-44e1-909f-ceca805a7b5d · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Training data-efficient image transformers & distillation through at- tention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.349483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.702916Z digest=sha256:1da856d9b7aabe58e6db606739aefa1c904d12c9c6e11a12f82a36f2197e126d

Observation e3b1bac5-b829-415c-8c37-3cb2e6eb3a82 · outbound

This paper cites Maxvit: Multi-axis vision transformer, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Maxvit: Multi-axis vision transformer, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.247621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.707587Z digest=sha256:fd410fd59dffc8edb7551246b031fc9ef8c26dcb4d5fad35ed33b6cfa58a4d9b

Observation a68175a4-aef5-436b-a78d-6f4fad36d37e · outbound

This paper cites In- stance normalization: The missing ingredient for fast styliza- tion, 2017.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers In- stance normalization: The missing ingredient for fast styliza- tion, 2017

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.233330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.712066Z digest=sha256:09ec8060da7ce75ff6bb3055915a6b03cc9e15e5f8b52206afa3b0b859f2edb7

Observation 010ca6d3-2c3b-44b4-ae3d-80234adaec9c · outbound

This paper cites Attention is all you need.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Attention is all you need

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.218969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.716541Z digest=sha256:0245a6128bc6ad750815c6d6eef2786935500df34ed7c5e158be0e462042cc18

Observation 7ccb88f5-0263-4385-bd4d-493a491ee74b · outbound

This paper cites When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.721130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.721130Z digest=sha256:7035dfa569f6544ef5cf0c0678a5e4d67d9dac8e18ba423a4b0ea95751aafff0

Observation 9d642eae-4c45-4e06-ae32-ec959ce8413e · outbound

This paper cites Ri- former: Keep your vision backbone effective but removing token mixer.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Ri- former: Keep your vision backbone effective but removing token mixer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.188014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.725421Z digest=sha256:91256c46dd2b16a020110691e30e7b03635fef0d3ae890af47d9619268ab0851

Observation 934a293f-64b3-41c7-9779-35f66fab10b0 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.066818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.729739Z digest=sha256:988a087298ca63c63a845816282dd790afacd81af40e7ec2570429ed81fb5a33

Observation 58281994-c54f-402d-868d-32042a40b407 · outbound

This paper cites PVT v2: Improved baselines with pyramid vision transformer.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers PVT v2: Improved baselines with pyramid vision transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.013124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.734132Z digest=sha256:5abed298559789202418846504d0bc91b770d32c6d079a24b3b8955ee17922d1

Observation 71d710a3-29b9-41b8-8a47-834aab5ee576 · outbound

This paper cites Active Token Mixer.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Active Token Mixer

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:43:32.104035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.748085Z digest=sha256:21c7ccfbdd72dbdbdf7d9e37586ed6a1cebc1d5aa50aa2f9b4ac878d546576ff

Observation aab8e676-b42c-4244-a709-ec76e82795de · outbound

This paper cites Con- vnext v2: Co-designing and scaling convnets with masked autoencoders.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Con- vnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.998901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.777059Z digest=sha256:0a43c24bee77e8d270cc3b14077b6ef74e4d27dcd06e91a763fdaa7a41fc7fce

Observation b7f03e5f-06f8-4ac8-a329-0c203d9ea072 · outbound

This paper cites Towards stabilizing batch statistics in backward propagation of batch normalization, 2020.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Towards stabilizing batch statistics in backward propagation of batch normalization, 2020

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.983278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.835347Z digest=sha256:81cb18624e05cc34989bedca1ef79dd372bf77b88319c79c05a0e94747bd6f5f

Observation fcb1aa8c-c73d-4f60-a7e9-adfe1931fdef · outbound

This paper cites Focal Self-attention for Local-Global Interactions in Vision Transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.912949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.912949Z digest=sha256:c37f3cdde290bc9590f9bcc98c1f3471831c4918d5a3bb5739dfad04945e7f82

Observation 60e22678-5467-4c24-bf79-d83fd66aa814 · outbound

This paper cites Focal modulation networks, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal modulation networks, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.915670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.940912Z digest=sha256:15ca1da57e77b69f1a153f21dcafae1905b2653fbef624b2ebdc09d4688cf8f4

Observation 2b092c62-72a5-4398-a493-7d7e1cf01390 · outbound

This paper cites Leveraging batch normalization for vision transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Leveraging batch normalization for vision transformers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.836864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.952058Z digest=sha256:d2f5ee9801869c185e1b2b9bab40c68c264c726a6781d5a373dc3f56dbffd5f7

Observation 482c18dc-8972-4260-b143-2d85abeea79a · outbound

This paper cites S2-mlp: Spatial-shift mlp architecture for vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers S2-mlp: Spatial-shift mlp architecture for vision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.687788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.956833Z digest=sha256:e229982da0ea2cac7fb25f61f850d93d630343e58a67fd17f685dfe212a2966b

Observation 3d9b8806-7a7e-4cbd-bc12-361009d8e795 · outbound

This paper cites Metaformer is actually what you need for vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Metaformer is actually what you need for vision

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.618259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.961673Z digest=sha256:76f2c3e87ad76c293d86f1d218e7a171433c2ff43cde25dfe8997eb904409443

Observation 377dfb90-3f52-4454-99f2-561b28605c99 · outbound

This paper cites Metaformer baselines for vision, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Metaformer baselines for vision, 2022

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.601707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.965697Z digest=sha256:1178799b436a7fd3ea2faf50050acfdfe5ed6f0e3d6ca4d8499c306a1e01de3f

Observation 686a3d8f-9eaf-4655-bd66-a73e7f99c996 · outbound

This paper cites Inceptionnext: When inception meets convnext, 2023.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Inceptionnext: When inception meets convnext, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.545384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.970308Z digest=sha256:877ce38e0779da21255e35315970c922c26f738bd40521137dda221df3f0f301

Observation fb307237-2f1d-4ee6-822f-a8b74b031bf0 · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.974627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.974627Z digest=sha256:2e73cc8c8a6efb3e822f00e26ab78f8c37aeb994ae09346bf1d02f9604b03e3b

Observation 0dd0aee0-f4f8-4da9-bff4-12b173bcb47f · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers mixup: Beyond Empirical Risk Minimization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.979103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.979103Z digest=sha256:0581b79d711c337b54f5df337a1ff534726e3a5bb424d04e8be79c1bebff7d96

Observation 0c849c23-b374-4e99-a96e-0a26c994c2d6 · outbound

This paper cites Random erasing data augmentation.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Random erasing data augmentation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.420065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.983603Z digest=sha256:91d4580e9a3caa8c3f01101ced8530e48995c0b3021e91807d124006225ee3ce

Observation efa4ee6a-aa2f-4b95-886f-71b7b68a6f3a · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Semantic under- standing of scenes through the ade20k dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.346275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T10:43:31.987951Z digest=sha256:9acc14d9821913454fd8f207fcda32e092495762f6c063ee2234ae7db8a78134

Pith citing papers

No inbound Pith citation observations are available.