REVIEW 3 major objections 6 minor 55 references
Vertical Fusion: Condensing Internal Representations for Robust ViT Classification
T0 review · 3 major / 6 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Intermediate Vision Transformer layers correct 18–76% of last-layer mistakes, and a learned vertical fusion of those layers closes 45% of the gap to an any-layer oracle.
desk verdict Solid multi-dataset recoverability numbers and a practical vertical fusion head; the diversity-vs-redundancy story is correlational, not causal, but the accuracy claims hold without it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
VFusion: a learnable fusion encoder that concatenates selected intermediate [CLS] tokens, maps the concatenation into a low-dimensional latent (with optional orthogonalization), and trains a lightweight classifier on that latent while the backbone remains frozen; together with the recoverability rate (fraction of final-layer errors corrected by at least one intermediate probe).
What would settle it
On a held-out collection of classification datasets, measure layer-wise recovery rates and train VFusion against best-layer and aggregation baselines; if recovery rates collapse toward zero and VFusion no longer beats the best single layer or the oracle gap remains largely unclosed, the central claim fails.
Extended reading notes
Core claim
Intermediate representations inside a single frozen Vision Transformer contain substantial corrective signal: independent layer-wise probes recover 18–76% of last-layer errors across 16 datasets. That recoverability is better explained by distributed, redundant probes of a shared decision boundary than by ensemble-style disagreement. A lightweight supervised fusion head (VFusion) that compresses the multi-layer hierarchy into one low-dimensional token captures a large fraction of this unused signal, closing 45% of the accuracy gap between the best single layer and a theoretical any-layer oracle, and outperforming established aggregation methods in both in-distribution and out-of-distribution
Load-bearing premise
That the corrective power of intermediate layers comes mainly from them being stable, redundant probes of one shared signal rather than from genuine predictive diversity, a claim resting on correlations with correction entropy and logit disagreement rather than a direct causal test.
Editorial extensions
If this is right
- A frozen off-the-shelf ViT can be made more accurate by attaching only a small fusion head instead of training or running multiple backbones.
- Horizontal ensembles become less necessary when the vertical hierarchy inside one model already supplies recoverable signal.
- Deeper contiguous layers preserve most of the fusion benefit, so practitioners can safely drop early layers if compute is tight.
- The same compression principle extends to horizontal fusion across heterogeneous pre-trained backbones when multiple models are available.
- Gains appear larger on fine-grained and degraded-image tasks, suggesting the method is most useful where last-layer features are fragile.
Reading between the lines
- The same recoverability pattern may appear in other transformer stacks (language, multimodal) that keep relatively uniform representations across depth, inviting vertical fusion heads beyond vision classification.
- If the redundancy-correctness account holds, unsupervised or self-supervised pre-training of the fusion encoder could remove the need for task labels at fusion time.
- Dense prediction tasks such as segmentation may benefit even more, because intermediate layers already encode spatial detail that the final [CLS] token discards.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that intermediate Vision Transformer layers contain substantial corrective signal for multi-class classification: independent layer-wise probes recover 18%–76% of last-layer errors across 16 datasets (Table 1; recovery rate in Eq. 2). It attributes this primarily to a “redundancy-correctness correspondence” rather than predictive diversity, based on correlations of recovery rate with correction entropy and (negatively) with pairwise logit disagreement (Fig. 1; §3.2). Building on that analysis, it proposes VFusion: select multi-layer [CLS] tokens, concatenate them, and compress them with a supervised MLP encoder into a low-dimensional latent (Eqs. 6–7) trained with cross-entropy plus an orthogonalization regularizer (Eqs. 8–10). Empirically, VFusion outperforms non-parametric and parametric aggregation baselines (Average, Majority, Super Learner, NLC, MoE, sw-MoE) on 16 ID datasets (Table 2; avg. 91.4% vs best-layer 89.5%, closing ~45% of the best-layer–oracle gap) and five shift benchmarks (Table 3), with ablations on layer selection, backbone family/scale, horizontal fusion, capacity, efficiency, and latent size.
Significance. If the empirical results hold, the work is a useful and practical contribution to frozen-backbone ViT deployment: it quantifies internal recoverability at scale, shows that simple last-layer probing leaves unused signal, and provides a lightweight vertical fusion head that improves ID and OOD accuracy without multi-backbone ensembles. Strengths include broad public-benchmark evaluation (21 datasets), multiple established baselines, capacity-controlled Final-only comparison (Appendix C.2), efficiency vs horizontal Super Learner (Appendix C.1), generalization across supervised/self-supervised/CLIP backbones and scales (Fig. 3), and released code. The recoverability framing and layer-selection guideline (contiguous depth preferred over striding) are actionable for practitioners. The mechanistic “redundancy-correctness” story is softer than the accuracy claims and should be treated as interpretive rather than established; even so, the method and measurement contributions remain significant for the field.
major comments (3)
- [§3.2, Fig. 1, Abstract] §3.2 and Fig. 1 (also Abstract/Introduction): the claim that recoverability is “not primarily driven by predictive diversity, but by a redundancy-correctness correspondence” is supported only by observational correlations (recovery rate vs H_corr, r=0.87; vs average pairwise cosine distance on jointly misclassified samples, r=−0.77; JS check in Appendix A.1). These do not causally isolate shared-signal compression from residual local complementarity (Appendix A.2 documents non-zero unique neighbor recoveries at every depth). Please rephrase causal language to correlational evidence, and clarify that VFusion’s compression design is motivated by—not proven by—this analysis. The accuracy claims do not require a stronger mechanism proof, but the current wording overstates what Fig. 1 establishes.
- [§4.2.1, Eqs. (6)–(7)] §4.2.1 / Eqs. (6)–(7): the decomposition h^(ℓ)=s^(ℓ)+r^(ℓ) and the assertion that the low-dimensional encoder “discards layer-specific residuals” are motivational, not measured. No ablation isolates shared vs layer-private components (e.g., reconstruction of s, mutual information across layers, or comparison of low-dim fusion vs a capacity-matched high-dim concat MLP without aggressive bottleneck). Appendix C.2 shows multi-layer input helps vs Final-only, and MoE uses the same concat features, but neither confirms noise-suppression of r^(ℓ). Either add a targeted diagnostic or present the encoder as supervised multi-layer dimensionality reduction without claiming residual filtering as established.
- [Table 2, §5.2, Eq. (1)] Table 2 / §5.2: VFusion approaches or exceeds the layer-wise oracle on Cars, Flow, IN1k, and ESAT. The oracle (Eq. 1) marks a sample correct if any single-layer probe is correct; surpassing it implies the fused representation creates new decision boundaries, not only selection among layers. This is interesting but under-discussed. Please quantify how often VFusion is correct when all layer probes fail (or when only a minority succeed), and state clearly that the oracle is not an upper bound on feature-level fusion—only on discrete layer selection—so the “45% of headroom closed” framing does not imply proximity to an absolute ceiling.
minor comments (6)
- [Tables 1–2] Tables consistently typeset “ESA T” / “ESA T” with a space (EuroSAT). Fix throughout Tables 1–2 and related captions.
- [§5.1.3] §5.1.3: probe hyperparameter grid and VFusion settings are clear; please also state whether MoE/sw-MoE expert MLPs match the probe head capacity and whether routing is trained jointly with experts under the same early-stopping protocol.
- [Figure 2] Figure 2 right panel is effective; ensure the “45%” annotation is defined in the caption as (a_VFusion − a_best)/(a_oracle − a_best) using the Table 2 averages so the figure is self-contained.
- [§6.3, Table 5] §6.3 / Table 5: HFusion is introduced as horizontal feature fusion; a one-sentence architectural parallel to VFusion (same encoder family on concatenated last-layer features) would help readers map vertical vs horizontal settings.
- [Appendix A.2] Appendix A.2 Table A1 is informative; consider a brief pointer in main §3.2 that local unique recoveries exist at all depths, so redundancy is imperfect—this would balance the low-disagreement narrative without new experiments.
- [§7.1] Limitations (§7.1) correctly note supervised fusion and frozen backbones; a short remark that recoverability is measured with trained probes (not zero-shot) would prevent over-reading for CLIP-style settings already discussed in §2.3.
Circularity Check
No significant circularity: recoverability and VFusion gains are measured against external benchmarks and independent baselines, not forced by definition or self-citation.
full rationale
This is an empirical methods paper. Recoverability (Eq. 2) is defined as the fraction of last-layer errors corrected by any intermediate probe and is measured on held-out test sets across 16 public datasets (Table 1); the oracle (Eq. 1) is an upper bound, not a fitted target that VFusion is forced to match. VFusion is a learned encoder+head trained with cross-entropy plus an optional orthogonalization regularizer (Eqs. 7–10) on concatenated frozen [CLS] tokens; its reported accuracy (Tables 2–3) is compared to independent baselines (Best layer, Average, Majority, SL, NLC, MoE, sw-MoE) and a capacity-controlled Final-only ablation (Appendix C.2). Closing 45% of the best-layer-to-oracle gap is an empirical observation, not an algebraic identity. The redundancy-correctness story (Fig. 1 correlations of recovery rate with correction entropy and negative correlation with logit cosine/JS disagreement) is interpretive support for the design choice of compression over diversity; it does not define the accuracy metric or force the numbers. No uniqueness theorem, ansatz smuggled via self-citation, or fitted parameter renamed as prediction appears. Self-citations are ordinary related-work references and are not load-bearing for the central claims. Score 0 is appropriate.
Assumptions & free parameters
free parameters (4)
- latent dimension dz
- orthogonalization weight λ
- fusion encoder widths [2048,1024,512]
- probe and fusion optim hyperparameters
assumptions (4)
- domain assumption Layer-wise [CLS] tokens are adequate global descriptors of intermediate ViT representations for classification probes and fusion.
- domain assumption A frozen pretrained backbone plus a supervised lightweight head is a valid setting for measuring and exploiting internal corrective signal.
- ad hoc to paper h^(ℓ) = s^(ℓ) + r^(ℓ) with a low-dimensional encoder isolating shared task-relevant s across layers.
- ad hoc to paper Recovery rate and correction entropy quantify useful internal corrective capacity for multi-class classification.
invented entities (3)
-
recoverability / recovery rate
independent evidence
-
redundancy-correctness correspondence
-
VFusion fusion encoder + orthogonalized latent global token
independent evidence
Cite this review
Pith. "Pith review of Vertical Fusion: Condensing Internal Representations for Robust ViT Classification." pith.science (2026). https://pith.science/paper/F5FSVYIO
@misc{pith2026260710391,
author = {Pith},
title = {Pith review of: Vertical Fusion: Condensing Internal Representations for Robust ViT Classification},
year = {2026},
howpublished = {\url{https://pith.science/paper/F5FSVYIO}},
note = {Machine review of arXiv:2607.10391}
}
read the original abstract
Despite exposing rich intermediate representations, Vision Transformers (ViTs) are almost exclusively utilized as black-box feature extractors, where only the last layer is considered for downstream tasks. We challenge this convention by introducing the notion of recoverability: the capacity of intermediate representations to correct last-layer failures. By evaluating independent classification probes at every model depth across 16 datasets, we observe that intermediate probes correctly classify 18% to 76% of samples that the last-layer probe misclassifies. We show that these gains are not primarily driven by predictive diversity, but by a redundancy-correctness correspondence, where the internal hierarchy acts as a series of stable, redundant probes of a shared discriminative signal. While established horizontal ensemble strategies (i.e., across multiple models) can improve performance, they incur high computational cost and ignore this vertical signal within a single model. To bridge this gap, we propose VFusion, a principled vertical aggregation strategy employing a learnable mapping into a low-dimensional latent space that synthesizes features across the internal ViT hierarchy. VFusion substantially outperforms established aggregation baselines in both in-distribution and out-of-distribution settings, notably closing 45% of the accuracy gap between the best individual layer and a theoretical oracle performance. Our gains consistently generalize across model sizes and pre-training regimes, confirming that VFusion offers a robust and efficient alternative to horizontal ensemble methods. The code is available at https://github.com/francescodisalvo05/vit-vertical-fusion.
Figures
Reference graph
Works this paper leans on
-
[1]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Un- terthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, An image is worth 16x16 words: Transformers for image recog- nition at scale, in: International Conference on Learning Representations, 2021
2021
-
[2]
Ranftl, A
R. Ranftl, A. Bochkovskiy, V . Koltun, Vision transformers for dense prediction, in: Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 12179–12188. 23
2021
-
[3]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, P. Bojanowski, DINOv2: Learning robust visual features without super...
2024
-
[4]
Teterwak, K
P. Teterwak, K. Saito, T. Tsiligkaridis, B. A. Plummer, K. Saenko, Is large- scale pretraining the secret to good domain generalization?, in: The Thirteenth International Conference on Learning Representations, 2025
2025
-
[5]
R. Imam, R. Marew, M. Yaqub, On the robustness of medical vision-language models: Are they truly generalizable?, in: Annual Conference on Medical Image Understanding and Analysis, Springer, 2025, pp. 233–256
2025
-
[6]
Kumar, N
D. Kumar, N. Muhammad, Object detection in adverse weather for autonomous driving through data merging and yolov8, Sensors 23 (20) (2023) 8471
2023
-
[7]
W. He, Z. Jiang, T. Xiao, Z. Xu, Y . Li, A survey on uncertainty quantification methods for deep learning, ACM Computing Surveys 58 (7) (2025)
2025
-
[8]
Minderer, J
M. Minderer, J. Djolonga, R. Romijnders, F. Hubis, X. Zhai, N. Houlsby, D. Tran, M. Lucic, Revisiting the calibration of modern neural networks, Ad- vances in neural information processing systems 34 (2021) 15682–15694
2021
Show all 55 references
-
[9]
Pinto, P
F. Pinto, P. H. Torr, P. K. Dokania, An impartial take to the cnn vs transformer robustness contest, in: European conference on computer vision, Springer, 2022
2022
-
[10]
Gulrajani, D
I. Gulrajani, D. Lopez-Paz, In search of lost domain generalization, in: Interna- tional Conference on Learning Representations, 2021
2021
-
[11]
Hendrycks, T
D. Hendrycks, T. Dietterich, Benchmarking neural network robustness to com- mon corruptions and perturbations, Proceedings of the International Conference on Learning Representations (2019). 24
2019
-
[12]
Raghu, T
M. Raghu, T. Unterthiner, S. Kornblith, C. Zhang, A. Dosovitskiy, Do vision transformers see like convolutional neural networks?, Advances in neural infor- mation processing systems 34 (2021) 12116–12128
2021
-
[13]
Wei, B.-L
T. Wei, B.-L. Wang, J.-X. Shi, Y .-F. Li, M.-L. Zhang, X-mahalanobis: Trans- former feature mixing for reliable OOD detection, in: The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
2025
-
[14]
Jeleni ´c, J
F. Jeleni ´c, J. Juki ´c, M. Tutek, M. Puljiz, J. Snajder, Out-of-distribution detec- tion by leveraging between-layer transformation smoothness, in: The Twelfth International Conference on Learning Representations, 2024
2024
-
[15]
Rodriguez-Opazo, D
Imezadelajara, C. Rodriguez-Opazo, D. Teney, D. Ranasinghe, E. Abbasnejad, Mysteries of the deep: Role of intermediate representations in out of distribu- tion detection, in: The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
2025
-
[16]
Uselis, S
A. Uselis, S. J. Oh, Intermediate layer classifiers for OOD generalization, in: The Thirteenth International Conference on Learning Representations, 2025
2025
-
[17]
Rodriguez-Opazo, E
C. Rodriguez-Opazo, E. Abbasnejad, D. Teney, H. Damirchi, E. Marrese-Taylor, A. van den Hengel, Synergy and diversity in CLIP: Enhancing performance through adaptive backbone ensembling, in: The Thirteenth International Con- ference on Learning Representations, 2025
2025
-
[18]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision, in: International Conference on Machine Learning (ICML), PmLR, 2021, pp. 8748–8763
2021
-
[19]
Nakata, Y
K. Nakata, Y . Ng, D. Miyashita, A. Maki, Y .-C. Lin, J. Deguchi, Revisiting a knn-based image classification system with high-capacity storage, in: European conference on computer vision, Springer, 2022, pp. 457–474
2022
-
[20]
Di Salvo, S
F. Di Salvo, S. Doerrich, I. Rieger, C. Ledig, An embedding is worth a thousand noisy labels, Transactions on Machine Learning Research (2025). 25
2025
-
[21]
Y . Zhu, J. Zhang, A. Gangrade, C. Scott, Label noise: Ignorance is bliss, Ad- vances in Neural Information Processing Systems 37 (2024) 116575–116616
2024
-
[22]
Mayilvahanan, R
P. Mayilvahanan, R. S. Zimmermann, T. Wiedemer, E. Rusak, A. Juhos, M. Bethge, W. Brendel, In search of forgotten domain generalization, in: The Thirteenth International Conference on Learning Representations, 2025
2025
-
[23]
Hendrycks, K
D. Hendrycks, K. Gimpel, A baseline for detecting misclassified and out-of- distribution examples in neural networks, in: International Conference on Learn- ing Representations, 2017
2017
-
[24]
Y . Sun, Y . Ming, X. Zhu, Y . Li, Out-of-distribution detection with deep nearest neighbors, International Conference on Machine Learning (ICML) (2022)
2022
-
[25]
Koutlis, S
C. Koutlis, S. Papadopoulos, Leveraging representations from intermediate encoder-blocks for synthetic image detection, in: European Conference on Com- puter Vision, Springer, 2024, pp. 394–411
2024
-
[26]
T. G. Dietterich, Ensemble methods in machine learning, in: International work- shop on multiple classifier systems (MCS), Springer, 2000, pp. 1–15
2000
-
[27]
Freund, R
Y . Freund, R. E. Schapire, A decision-theoretic generalization of on-line learning and an application to boosting, Journal of computer and system sciences 55 (1) (1997) 119–139
1997
-
[28]
Lakshminarayanan, A
B. Lakshminarayanan, A. Pritzel, C. Blundell, Simple and scalable predictive uncertainty estimation using deep ensembles, Advances in neural information processing systems 30 (2017)
2017
-
[29]
S. Fort, H. Hu, B. Lakshminarayanan, Deep ensembles: A loss landscape per- spective, arXiv preprint arXiv:1912.02757 (2019)
1912 arXiv
-
[30]
Wortsman, G
M. Wortsman, G. Ilharco, S. Y . Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y . Carmon, S. Kornblith, et al., Model soups: averaging weights of multiple fine-tuned models improves accuracy with- out increasing inference time, in: International c...
2022
-
[31]
Ainsworth, J
S. Ainsworth, J. Hayase, S. Srinivasa, Git re-basin: Merging models modulo permutation symmetries, in: The Eleventh International Conference on Learning Representations, 2023
2023
-
[32]
K. Lenc, A. Vedaldi, Understanding image representations by measuring their equivariance and equivalence, in: Proceedings of the IEEE conference on com- puter vision and pattern recognition, 2015, pp. 991–999
2015
-
[33]
Bansal, P
Y . Bansal, P. Nakkiran, B. Barak, Revisiting model stitching to compare neural representations, Advances in neural information processing systems 34 (2021)
2021
-
[34]
C. Ju, A. Bibaut, M. van der Laan, The relative performance of ensemble meth- ods with deep convolutional neural networks for image classification, Journal of applied statistics 45 (15) (2018) 2800–2818
2018
-
[35]
Zbontar, L
J. Zbontar, L. Jing, I. Misra, Y . LeCun, S. Deny, Barlow twins: Self-supervised learning via redundancy reduction, in: International conference on machine learning, PMLR, 2021, pp. 12310–12320
2021
-
[36]
Bardes, J
A. Bardes, J. Ponce, Y . LeCun, VICReg: Variance-invariance-covariance regu- larization for self-supervised learning, in: International Conference on Learning Representations, 2022
2022
-
[37]
Krizhevsky, G
A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images (2009)
2009
-
[38]
Fei-Fei, R
L. Fei-Fei, R. Fergus, P. Perona, Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object cate- gories, Computer vision and Image understanding 106 (1) (2007) 59–70
2007
-
[39]
Krause, M
J. Krause, M. Stark, J. Deng, L. Fei-Fei, 3d object representations for fine- grained categorization, in: Proceedings of the IEEE international conference on computer vision workshops, 2013, pp. 554–561
2013
-
[40]
C. Wah, S. Branson, P. Welinder, P. Perona, S. Belongie, The caltech-ucsd birds- 200-2011 dataset (2011). 27
2011
-
[41]
Cimpoi, S
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, A. Vedaldi, Describing textures in the wild, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 3606–3613
2014
-
[42]
Helber, B
P. Helber, B. Bischke, A. Dengel, D. Borth, Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 12 (7) (2019) 2217–2226
2019
-
[43]
S. Maji, E. Rahtu, J. Kannala, M. Blaschko, A. Vedaldi, Fine-grained visual classification of aircraft, arXiv preprint arXiv:1306.5151 (2013)
2013 arXiv
-
[44]
Nilsback, A
M.-E. Nilsback, A. Zisserman, Automated flower classification over a large num- ber of classes, in: Indian Conference on Computer Vision, Graphics and Image Processing, 2008
2008
-
[45]
Bossard, M
L. Bossard, M. Guillaumin, L. Van Gool, Food-101 – mining discriminative components with random forests, in: European Conference on Computer Vision, 2014
2014
-
[46]
Houben, J
S. Houben, J. Stallkamp, J. Salmen, M. Schlipsing, C. Igel, Detection of traffic signs in real-world images: The German Traffic Sign Detection Benchmark, in: International Joint Conference on Neural Networks, no. 1288, 2013
2013
-
[47]
LeCun, C
Y . LeCun, C. Cortes, C. Burges, Mnist handwritten digit database, ATT Labs [Online]. Available: http://yann.lecun.com/exdb/mnist 2 (2010)
2010
-
[48]
B. S. Veeling, J. Linmans, J. Winkens, T. Cohen, M. Welling, Rotation equiv- ariant cnns for digital pathology, in: International Conference on Medical image computing and computer-assisted intervention, Springer, 2018, pp. 210–218
2018
-
[49]
O. M. Parkhi, A. Vedaldi, A. Zisserman, C. Jawahar, Cats and dogs, in: 2012 IEEE conference on computer vision and pattern recognition, IEEE, 2012, pp. 3498–3505. 28
2012
-
[50]
Coates, A
A. Coates, A. Ng, H. Lee, An analysis of single-layer networks in unsupervised feature learning, in: Proceedings of the fourteenth international conference on ar- tificial intelligence and statistics, Journal of Machine Learning Research (JMLR) Workshop and Conference Proceedi...
2011
-
[51]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255.doi:10.1109/CVPR.2009.5206848
2009 doi
-
[52]
Z. Liu, P. Luo, X. Wang, X. Tang, Deep learning face attributes in the wild, in: Proceedings of the IEEE international conference on computer vision, 2015, pp. 3730–3738
2015
-
[53]
Sagawa, P
S. Sagawa, P. W. Koh, T. B. Hashimoto, P. Liang, Distributionally robust neural networks, in: International Conference on Learning Representations, 2020
2020
-
[54]
Shazeer, A
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, J. Dean, Out- rageously large neural networks: The sparsely-gated mixture-of-experts layer, arXiv preprint arXiv:1701.06538 (2017)
2017 arXiv
-
[55]
Fedus, B
W. Fedus, B. Zoph, N. Shazeer, Switch transformers: Scaling to trillion pa- rameter models with simple and efficient sparsity, Journal of Machine Learning Research 23 (120) (2022) 1–39. 29 Vertical Fusion: Condensing Internal Representations for Robust ViT Classification Appen...
2022
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.