Pith. sign in

Paper Citation Record · LEDGER

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers

As of 13 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2411.18995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18995 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:43:31.987951Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c7de7d48-ba92-4b66-8577-9f33223c599b · outbound

This paper cites Layer Normalization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.467766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.467766Z digest=sha256:522c83afc718c880b675c8931edc2515392548d2e0e27a7e9c6efd1a3b3b081c

Observation 7475193f-5158-467a-ace5-713f4e0b2708 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.508733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.508733Z digest=sha256:377fbabc23bd1f858649d149a593171b5ba168ad6ccd55ecaaf6bc55fd2341b1

Observation 10321d41-0599-4d38-80f6-c09ac586fcb8 · outbound

This paper cites CycleMLP: A MLP-like Architecture for Dense Prediction.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers CycleMLP: A MLP-like Architecture for Dense Prediction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.546956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.546956Z digest=sha256:131e4197a14d41b939a6093eb8ee6d831f6c30994f94a25a27e2cce63a6ea9c3

Observation 39a28443-e4a1-4509-9530-e1271565b2c6 · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers, 2021.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Twins: Revisiting the design of spatial attention in vision transformers, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:35.070608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.552877Z digest=sha256:51e405c373f7954be919bfac5badec9ea241c6e732d73e83c70f3335cf23b291

Observation 38fee022-f3ce-46fc-9af3-cb39d6b99425 · outbound

This paper cites Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.627245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.627245Z digest=sha256:063f39c20115fd3411f645d7af29a94d6122392e219af6a90ec7ebe117396c36

Observation d348127d-e7fc-4b8a-84af-5b9fa2414222 · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.917980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.688855Z digest=sha256:64ba468b33bb08edccdd401df82aa50c439e8309a1cfd43aa7122ba2c7592827

Observation aaa21eac-edf0-464d-87df-3800442065ce · outbound

This paper cites Coatnet: Marrying convolution and attention for all data sizes.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Coatnet: Marrying convolution and attention for all data sizes

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.859205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.738670Z digest=sha256:13f59fddfd3e35396b7034787aa6a894d735cab9a7c1dcdb4b983b10f742a10c

Observation 2c4525cc-c989-4802-9513-0e2f03475a57 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Imagenet: A large-scale hierarchical image database

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.832140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.744332Z digest=sha256:2c7fdde91487d695cc0e387cf560cfe0fad70944ed13e1e49e7c5f30c0afdf8e

Observation 9841c4be-b33b-4317-beb2-93eade38f88f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.796276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.796276Z digest=sha256:d5343df092d22ccb169dd0db93637a27b82d267bd6c8dc374222b8e2b8737bf2

Observation f0d87a41-ce01-46e2-bcab-4827e813cb0d · outbound

This paper cites Multiscale Vision Transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.849424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.849424Z digest=sha256:f122bf766796c1e78c18311166d68c50ba6929a0c2096e7fa749577f79b70906

Observation bd35d36f-b548-49eb-af81-00f068998dad · outbound

This paper cites Segnext: Rethinking convolutional attention design for semantic segmentation,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Segnext: Rethinking convolutional attention design for semantic segmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.726346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.872013Z digest=sha256:40cc160b72a497b630dcfb783a1ac2cdbbb8838f0a50d38bbc4d5e58097252c0

Observation e9170c0f-c16c-4480-8fb4-a70225783446 · outbound

This paper cites Visual Attention Network.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Visual Attention Network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:30.877886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:30.877886Z digest=sha256:5da89645dabd02c570950d50db7bbc1728afa3c89ed17e831ff887d0b0e7ed16

Observation b4b86f3a-ccc3-42c6-9343-2aaeb8f003ab · outbound

This paper cites Deep residual learning for image recognition.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Deep residual learning for image recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.617654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.900984Z digest=sha256:fe90091e799eb923f479c11a99d0ab200338542bdbaef3fdeb35afbde8882304

Observation 1b8d0802-2d77-4adb-93bf-6fad9df567e1 · outbound

This paper cites Mask r-cnn, 2018.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mask r-cnn, 2018

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.602047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.967885Z digest=sha256:509225af246c2f18cdbd72ea864a139d0a808c6b8ce1f10634e5198d6a210977

Observation 9a869c6b-0f24-41a6-b7d7-23d3e7b6be31 · outbound

This paper cites Deep networks with stochastic depth.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Deep networks with stochastic depth

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.556396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:30.997621Z digest=sha256:eff2b9e80ae15bf24889414f291f1594e8ddb89e749c6e2b95e0f2a4d7d614d9

Observation bed20438-0074-4926-b6e5-a5a779e30582 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization, 2017.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Arbitrary style transfer in real-time with adaptive instance normalization, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.424854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.038529Z digest=sha256:98488de94678910c4a7b74a8417a78582f13056da8dce999e9696eba8f1bb13b

Observation 29986528-00a0-4f41-93d0-31d13389fea1 · outbound

This paper cites Batch renormalization: Towards reducing minibatch dependence in batch-normalized models, 2017.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Batch renormalization: Towards reducing minibatch dependence in batch-normalized models, 2017

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.356058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.051916Z digest=sha256:a200f6f03af07024d75271556c5bab2809f54e17bbcb0f16f62228d727df9064

Observation e595f8bc-5db7-46cf-a31e-7b9e1280c0fe · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.056554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.056554Z digest=sha256:5a88019f0c71fdc798f43e41e9e8470c9f3e42234cc97eae9a1604c76470c064

Observation 1fa7156f-bf36-4b4d-8030-57372fbb9e40 · outbound

This paper cites Relational self-attention: What’s missing in attention for video understanding.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Relational self-attention: What’s missing in attention for video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.341086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.061886Z digest=sha256:47b1560ed366290b05c652ddbb3c2d4d6bbadb40afcd80739b5e285525afb6a2

Observation f2835c07-1940-4809-b32c-329e8f100596 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.086558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.086558Z digest=sha256:b7b2c8a7c947f83ec10f281ae1375e26d464d69057d8b30853d7f3304d928d39

Observation 5722a313-f18f-455e-935c-0cd05c2e3472 · outbound

This paper cites Panoptic feature pyramid networks.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Panoptic feature pyramid networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.308883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.148497Z digest=sha256:1f1045d5d46ceff525b2ed9d44f82f65c166c3445248209f463d3b9026ea7fc5

Observation a9de85a0-e180-435f-98e7-8c8daa20cef3 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mvitv2: Improved multiscale vision transformers for classification and detection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.154322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.193362Z digest=sha256:a69abb7c646b2d8f7bfbd73488b62c2676ba0d720b1b61e9a5d82a65106f3952

Observation d7166ac2-e40f-425b-b744-885ac63c940e · outbound

This paper cites AS-MLP: An Axial Shifted MLP Architecture for Vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers AS-MLP: An Axial Shifted MLP Architecture for Vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.249528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.249528Z digest=sha256:e1f7a9abd4db80c1f6573c14e999b7de5b2aeceacc0644d2ae37ffe59978f823

Observation 3ca758f5-1059-4aae-9d56-02f41a15dc4a · outbound

This paper cites Microsoft coco: Common objects in context.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Microsoft coco: Common objects in context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.273700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.273700Z digest=sha256:9e7a6667f59a5f8cde0ff9ab69b1dee400cd4132ba57ddfe269ff3d855f34446

Observation 331bfe55-e673-4641-87c1-7b13291d4340 · outbound

This paper cites Focal loss for dense object detection, 2018.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal loss for dense object detection, 2018

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.080426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.278519Z digest=sha256:27e423d0ca8d0d58aef422255b1c40039a0d496cd8fc0f445db40b7ba2dcdfc9

Observation 50af3cc3-dbaa-4b87-ae15-d1ae37dfcb64 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:34.057496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.283982Z digest=sha256:5d1faa2af860b78e06075bd330cfdadeed2323c9d1c5e41b567090b789fcd691

Observation 5732a23a-d106-486c-afc5-4c97c07cb58a · outbound

This paper cites A convnet for the 2020s.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers A convnet for the 2020s

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.288580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.288580Z digest=sha256:478bf0e0d23dafa2168d5ebdb4a05575a79e643b4733af537a0cf87767aa8112

Observation 11d4a1a3-db15-4655-913e-64721d343196 · outbound

This paper cites Decoupled Weight Decay Regularization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.345474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.345474Z digest=sha256:2d621d3bf551975aea5c8cf20d3ba17c1d1e4dc06018266eb7eda1cdf6fc47b1

Observation 3637f4e0-8dfe-4153-86c9-994b75885459 · outbound

This paper cites How do vision transformers work?, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers How do vision transformers work?, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.920037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.408851Z digest=sha256:c62c0ce1de1015f55961f64efff996a1855e16790c2da196ee0cdd2fab71faf9

Observation 66a09598-7ecb-4a39-b708-bd37f83e4407 · outbound

This paper cites Semantic image synthesis with spatially-adaptive nor- malization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Semantic image synthesis with spatially-adaptive nor- malization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.436305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.436305Z digest=sha256:feb56adc698ff68114121c4737ed92c1e03df4a702d72ded08cd5898c263e692

Observation 1d79b78b-fc4a-46b3-94eb-64f62fc40bc7 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Pytorch: An im- perative style, high-performance deep learning library

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.828294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.484874Z digest=sha256:125a6afd2daea03599078d252ded61c60ccf6b55d7fe4d93e80bad875d2723ed

Observation f007f04f-8fa1-4923-a4a2-595b4b2d6a9c · outbound

This paper cites Acceleration of stochastic approximation by averaging.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Acceleration of stochastic approximation by averaging

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.813907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.489768Z digest=sha256:c4c595ec900ea7e209178852131f10842e2d4ced59f3b5429b886aa0cafb2509

Observation 7ec5e319-f1ad-4b0e-8730-94dc35119152 · outbound

This paper cites What makes for good tokenizers in vision transformer? IEEE Transactions on Pattern Analysis and Machine Intelligence,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers What makes for good tokenizers in vision transformer? IEEE Transactions on Pattern Analysis and Machine Intelligence,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.799813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.494333Z digest=sha256:b7afc4d0724b7aed2e3bd6ef39e579aca71e6f0b17685ecf93e0669374c80d06

Observation 85eb279b-be1e-4cd5-971b-1d3f423a896c · outbound

This paper cites Do vision trans- formers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34:12116–12128,.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Do vision trans- formers see like convolutional neural networks? Advances in Neural Information Processing Systems, 34:12116–12128,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.499973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.499973Z digest=sha256:141930a701c3b1fc78d046f0c25e28e53c47c5943090ee6e1ccd45610729f5ae

Observation a066e8f0-2382-4d92-9860-4fdfb750d206 · outbound

This paper cites Mobilenetv2: Inverted residuals and linear bottlenecks, 2019.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mobilenetv2: Inverted residuals and linear bottlenecks, 2019

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.651012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.504794Z digest=sha256:6c5bf32563c3de476ade1ed2a41db35129153ab53a2ce89253d3a7369bc643a3

Observation 43ce6aa9-f109-4aee-8aa8-3c833791980f · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.561151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.509672Z digest=sha256:201d1d7719afa6dfaf454b3da5290682a8d0269fa0cadc55cfb92287197acb6c

Observation 00b7bd86-6602-4cb8-b4ae-862074991779 · outbound

This paper cites Powernorm: Rethinking batch normaliza- tion in transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Powernorm: Rethinking batch normaliza- tion in transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.546723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.514224Z digest=sha256:b3dc6047ac8676059ddacbf08fd204f33bd0d72b176ab08c6a2207a8dea51d04

Observation 5b13d27e-365c-4453-a5ce-616fe2a4970d · outbound

This paper cites NormFormer: Improved Transformer Pretraining with Extra Normalization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers NormFormer: Improved Transformer Pretraining with Extra Normalization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.585543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.585543Z digest=sha256:f4e878204691a9d5da339f78795624ac11f26cec8eb62b744934c43eef0bef3a

Observation b5ae794c-f344-4b07-9303-94610bc313a6 · outbound

This paper cites Evalnorm: Esti- mating batch normalization statistics for evaluation, 2019.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Evalnorm: Esti- mating batch normalization statistics for evaluation, 2019

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.532254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.661340Z digest=sha256:9f11b3946a46365492a235356a378d972a95331666361410a25c9fce4458d0ca

Observation 2dd01271-4d17-4d0b-a9c3-50a8174ddd71 · outbound

This paper cites Rethinking the inception archi- tecture for computer vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Rethinking the inception archi- tecture for computer vision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.689239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.689239Z digest=sha256:84cfbc496fd1abccd2a123988434c505211ae7e7c18ff645473463fdff55e0dd

Observation af4e1c54-1194-4a2e-a0ac-16d053fd602c · outbound

This paper cites Mlp-mixer: An all-mlp architecture for vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Mlp-mixer: An all-mlp architecture for vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.693505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.693505Z digest=sha256:66545bb9b2031984f1bdd81a75ceb1fe4dbef1a667a3fe606b2a9ffa88bb66d3

Observation 64fd3689-1a74-44e6-9569-c61fbf8cb174 · outbound

This paper cites Resmlp: Feedforward networks for image clas- sification with data-efficient training, 2021.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Resmlp: Feedforward networks for image clas- sification with data-efficient training, 2021

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.465429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.698267Z digest=sha256:db515dd5cf5e94ea6a3e165cdb69f364df47d4a7b1e150375093c66ead2d3f38

Observation a2ea657f-b050-44e1-909f-ceca805a7b5d · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Training data-efficient image transformers & distillation through at- tention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.349483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.702916Z digest=sha256:e0c55e43302ff2e18979924de84df56abdaa375f1bff2c76ff054b68fbea165e

Observation e3b1bac5-b829-415c-8c37-3cb2e6eb3a82 · outbound

This paper cites Maxvit: Multi-axis vision transformer, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Maxvit: Multi-axis vision transformer, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.247621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.707587Z digest=sha256:c00331f4b6eaeca7abb15191d5b86c4ff3e2bfd02f32d6a22550307f98dd2b95

Observation a68175a4-aef5-436b-a78d-6f4fad36d37e · outbound

This paper cites In- stance normalization: The missing ingredient for fast styliza- tion, 2017.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers In- stance normalization: The missing ingredient for fast styliza- tion, 2017

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.233330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.712066Z digest=sha256:40af23d8d42a10d8050b82a716576cd671f130f7a3c689549e89349281961760

Observation 010ca6d3-2c3b-44b4-ae3d-80234adaec9c · outbound

This paper cites Attention is all you need.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Attention is all you need

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.218969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.716541Z digest=sha256:8304311136e0eccede0656bbbd042cf657650016a70e23dc062831afe6180c28

Observation 7ccb88f5-0263-4385-bd4d-493a491ee74b · outbound

This paper cites When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.721130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.721130Z digest=sha256:32b54cb270b98e6a39944f600834ece141faff57a4308f3b708d6b0b635388ed

Observation 9d642eae-4c45-4e06-ae32-ec959ce8413e · outbound

This paper cites Ri- former: Keep your vision backbone effective but removing token mixer.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Ri- former: Keep your vision backbone effective but removing token mixer

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.188014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.725421Z digest=sha256:ec2e55624cad2ce2fca778e5d5b87fef95b0c88c429d7530aa9d5fa727177ec3

Observation 934a293f-64b3-41c7-9779-35f66fab10b0 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.066818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.729739Z digest=sha256:c040cb181285af2cbd9cab8242058ba324df1131466411abe43269de31e076ae

Observation 58281994-c54f-402d-868d-32042a40b407 · outbound

This paper cites PVT v2: Improved baselines with pyramid vision transformer.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers PVT v2: Improved baselines with pyramid vision transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:33.013124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.734132Z digest=sha256:94382920fad69fd60945b6185803ab88ef47132c397780e46e59bf71dc5de3fe

Observation 71d710a3-29b9-41b8-8a47-834aab5ee576 · outbound

This paper cites Active Token Mixer.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Active Token Mixer

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:43:32.104035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.748085Z digest=sha256:16ec8fb39cf950cc575c16fd15258730e36e90a0c1741801d0b3f17508e27372

Observation aab8e676-b42c-4244-a709-ec76e82795de · outbound

This paper cites Con- vnext v2: Co-designing and scaling convnets with masked autoencoders.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Con- vnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.998901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.777059Z digest=sha256:af86a8b61129683c31af586be6ca03b079a540fd165ea276aa7fb85ef2a40907

Observation b7f03e5f-06f8-4ac8-a329-0c203d9ea072 · outbound

This paper cites Towards stabilizing batch statistics in backward propagation of batch normalization, 2020.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Towards stabilizing batch statistics in backward propagation of batch normalization, 2020

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.983278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.835347Z digest=sha256:1144684ec2c411acccd0bffccebc70d86a0dc84196df0f74aae586ed00e7e9ab

Observation fcb1aa8c-c73d-4f60-a7e9-adfe1931fdef · outbound

This paper cites Focal Self-attention for Local-Global Interactions in Vision Transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.912949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.912949Z digest=sha256:d76e478e3d0c9ba145d063d3d90121c8378652b4f8cbf93dfa9477993c0fa108

Observation 60e22678-5467-4c24-bf79-d83fd66aa814 · outbound

This paper cites Focal modulation networks, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Focal modulation networks, 2022

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.915670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.940912Z digest=sha256:b9d8ff03dcb9216834beed674416c7a5fc729b079406522370123ac52f0e7e22

Observation 2b092c62-72a5-4398-a493-7d7e1cf01390 · outbound

This paper cites Leveraging batch normalization for vision transformers.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Leveraging batch normalization for vision transformers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.836864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.952058Z digest=sha256:b330709bcf8e5989539d2af524dfaa1c2eb5b2e419a6863e8aff9fe513377ddc

Observation 482c18dc-8972-4260-b143-2d85abeea79a · outbound

This paper cites S2-mlp: Spatial-shift mlp architecture for vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers S2-mlp: Spatial-shift mlp architecture for vision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.687788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.956833Z digest=sha256:f1e3f6fd2d668e9029de4d078ec799da378602c2a9b84830aa07810354a7c7cd

Observation 3d9b8806-7a7e-4cbd-bc12-361009d8e795 · outbound

This paper cites Metaformer is actually what you need for vision.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Metaformer is actually what you need for vision

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.618259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.961673Z digest=sha256:6d13d9583714b601634b45f7e11ca62a39c881c151f5bd0295acedbc04a28ea6

Observation 377dfb90-3f52-4454-99f2-561b28605c99 · outbound

This paper cites Metaformer baselines for vision, 2022.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Metaformer baselines for vision, 2022

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.601707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.965697Z digest=sha256:c821a1ccfdcfe86bdcbf35fc7d22953cc400390a7a65dbf67220843b7d07e6bf

Observation 686a3d8f-9eaf-4655-bd66-a73e7f99c996 · outbound

This paper cites Inceptionnext: When inception meets convnext, 2023.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Inceptionnext: When inception meets convnext, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.545384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.970308Z digest=sha256:40916fd5bf222468ffe7e58adc5f9e05fb6dd4ff22a458c28b4df6e906431261

Observation fb307237-2f1d-4ee6-822f-a8b74b031bf0 · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.974627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.974627Z digest=sha256:d9d8d98cfaf34db05bb8a2e216d9d8607524c35df0df45a001ce838f316dd2c1

Observation 0dd0aee0-f4f8-4da9-bff4-12b173bcb47f · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers mixup: Beyond Empirical Risk Minimization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T10:43:31.979103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:43:31.979103Z digest=sha256:a2ae8809c286605cc920e338a39ffdf5d090171a85d58cc7d761b7f50e64c4d8

Observation 0c849c23-b374-4e99-a96e-0a26c994c2d6 · outbound

This paper cites Random erasing data augmentation.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Random erasing data augmentation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.420065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.983603Z digest=sha256:4e895f627e035d52097ff8b4525452fe9b057e25dac514c8c0aa0f900ae01506

Observation efa4ee6a-aa2f-4b95-886f-71b7b68a6f3a · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers Semantic under- standing of scenes through the ade20k dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:43:32.346275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T10:43:31.987951Z digest=sha256:6a812fc46aa3598cac18fdf222d94ea46e75ffdec4401ac64a116847c699c9ef

Pith citing papers

No inbound Pith citation observations are available.