REVIEW 4 major objections 5 minor 77 references
Uncertainty Quantification for Incomplete Multi-View Data Using Divergence Measures
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Replacing KL divergence with Proper Hölder divergence tightens the evidence bound and improves multi-view learning under noise and missing data.
desk verdict An incremental but plausible evidential-fusion pipeline whose headline theoretical guarantee (PHD gives a tighter ELBO than KL) is asserted, not proven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Proper Hölder divergence (PHD) between two Dirichlet distributions, computed in closed form through the log-normalizer of the exponential family (Eq. (12)). This closed form makes the PHD-regularized ELBO trainable by gradient descent. Around it, the method wraps a variational Dirichlet layer that outputs evidence for each view, Dempster-Shafer combination rules to fuse modalities, and a Kalman filter that smooths the fused belief and predicts future states. The claim that PHD yields a tighter ELBO than KL is what converts the divergence swap into an accuracy and robustness gain.
What would settle it
Compute the PHD between two Dirichlet distributions using direct numerical integration of Definition 1 and compare it with the closed form in Eq. (12); a mismatch would falsify Theorem 1 and make the training loss in Eqs. (7)-(9) non-evaluable as written.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that the Proper Hölder divergence—defined through conjugate exponents and a scale exponent—admits a closed form for Dirichlet distributions, and that using this divergence as the regularization term in a variational objective yields a tighter ELBO than the KL divergence. Because the Dirichlet distribution models class probabilities as evidence, the tighter bound translates into better uncertainty estimates for each modality and for the fused decision. The authors further show that combining Dempster-Shafer evidence fusion with a Kalman filter improves robustness to missing views and future state estimation. The reported experiments support the claim that PHD-based training outperforms KL-based training on multi-view classification and clustering under noise and missing data.
Load-bearing premise
The entire training procedure assumes the closed-form expression for PHD between two Dirichlet distributions in Theorem 1 is algebraically correct; if the gamma-function expansions behind it are wrong, the loss function cannot be evaluated as written.
Editorial extensions
If this is right
- Any evidential deep learning pipeline that currently uses KL divergence in its Dirichlet prior can swap in PHD and inherit a tighter ELBO without changing the network architecture.
- Multi-view models will keep working when entire views are missing, because PHD-based training produces more reliable per-view uncertainties that Dempster-Shafer fusion can weight correctly.
- The Kalman-filtered fusion step gives a natural way to update beliefs online as new observations arrive, extending beyond the static fusion used in the ETMC baseline.
- Because PHD has tunable exponents $\alpha$, $\beta$, $\gamma$, the regularization strength can be adapted per dataset, which the grid-search results indicate is necessary for best performance.
Reading between the lines
- The closed-form PHD for Dirichlet distributions also applies to other exponential families with conic or affine natural parameter spaces, so the tighter-ELBO argument could be tested in Gaussian or categorical variational autoencoders.
- The ablation shows that adding the Dirichlet prior to Hölder divergence helps on ADE20K but slightly hurts on NYUD2 and SUN RGB-D, suggesting the practical recipe may need a data-dependent choice of whether to include the prior.
- A natural test is whether the reported accuracy gains persist when the Kalman filter and Dempster-Shafer fusion are ablated, isolating the contribution of PHD alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes KPHD-Net, a multi-view classification and clustering method that replaces the KL-divergence regularizer in an evidential Dirichlet variational objective with the Proper Hölder Divergence (PHD), fuses view-specific evidence with Dempster-Shafer theory, and adds a Kalman filter for dynamic state estimation. The authors claim a theoretical guarantee (Theorems 2–3) that PHD yields a tighter ELBO than KL in Dirichlet models, and support the method with experiments on four classification and five clustering datasets under noise and missing-data conditions. The appendix derives a closed-form PHD for Dirichlet distributions and includes a separate mode-collapse argument (Theorem 4).
Significance. If the central claim were established, the paper would offer a principled alternative to KL-based evidential fusion and a closed-form objective that is easy to optimize; the empirical scope is broad, covering classification, clustering, noise, missing rates, and multiple backbones, and the use of t-SNE visualizations is a plus. However, the advertised theoretical guarantee is not proven, the reported gains are entangled with per-dataset hyperparameter selection, and the ablation study partially contradicts the proposed design. As submitted, the contribution is therefore an empirical method with an unsubstantiated theoretical framing rather than a supported advance.
major comments (4)
- [Appendix A, Theorems 2 and 3] Theorems 2 and 3 are duplicate statements, and the proof does not prove the claim. After writing the KL and PHD expressions in Eqs. (19)–(22), the proof asserts that because PHD has tunable parameters α, β, γ it "can better fit the true posterior distribution and reduce the gap"; no inequality of the form D_PHD(q∥p) ≤ D_KL(q∥p) or any comparison of the full ELBOs is derived. Since the loss in Eqs. (7)–(9) replaces the KL regularizer with PHD while leaving the likelihood term unchanged, the claimed ELBO improvement requires exactly such an inequality at the hyperparameter values used. PHD is a projective divergence and is not generally ordered with respect to KL; the manuscript provides no argument for the Dirichlet case, so the central theoretical contribution is unsupported. The verbatim duplication of Theorem 3 should also be removed.
- [Section IV-I and Tables III, V, IX] The reported superiority is partly an artifact of post hoc hyperparameter selection. Table III lists up to five (α, γ) configurations per classification dataset, and Tables V and IX scan α and γ for Caltech-101-20 and ADE20K; the text then reports the best configuration as "Ours". No validation-based selection rule, number of random seeds, or performance at a fixed configuration is given, so the comparison with ETMC conflates the method with the grid search. The grid-search section even recommends different optimal α ranges for different datasets, reinforcing that the headline numbers are selected, not predictive. Please report a single default configuration or a nested validation procedure.
- [Table VIII (Ablation Study)] The ablation study undermines the claim that the full KPHD-Net design, including the Dirichlet model, is responsible for the gains. On NYUD2, Hölder without Dirichlet achieves 73.64% fusion accuracy versus 70.33% for Hölder with Dirichlet, and on SUN RGB-D the corresponding numbers are 61.60% versus 61.10%. The text attributes this to "marginal trade-offs", but it means the best reported fusion results for these datasets come from a component the paper does not propose as the final model. The authors need to either reconcile this with the method definition or revise the claim that the Dirichlet-based objective is the source of the improvement.
- [Theorem 4 (Appendix A)] The mode-collapse argument is not a proof. The proof states a Gaussian-mixture example and asserts that adjusting α, β, γ allows q to capture both modes, but it gives no calculation, no condition on α, β, γ, and no relation to the Dirichlet setting used in the paper. As with Theorems 2–3, this is an assertion about flexibility rather than a derivation, and it cannot support the claimed theoretical guarantees.
minor comments (5)
- [Section IV-A (SUNRGBD)] The dataset description reports both "4,845 training and 4,659 testing" and "20,210 and 3,000 samples" without explanation; this should be corrected.
- [Section IV-F (NYUD2 discussion)] The text says the model "slightly falls behind CNN-randRNN in the fusion modality (73.64% vs. 73.00%)" on NYUD2, but 73.64 is higher than 73.00; the sentence contradicts Table IV.
- [Equation (9)] Equation (9) writes PHD[D(µ_m|α_m)∥D(µ_m|[1,…,1])] with inconsistent notation compared with D_PHD used in Eqs. (7)–(8); the notation should be unified.
- [Abstract and Section I] The abstract says eight datasets while the text says four classification and five clustering datasets; the count should be reconciled.
- [Section IV-D] No code or implementation details (learning rate schedule, training epochs, number of trials) are provided, so the experiments are not reproducible as described.
Circularity Check
The advertised theoretical guarantee (Theorem 2, duplicated as Theorem 3) is a circular proof: it asserts ELBO_H ≥ ELBO_KL and then justifies it by saying PHD is tunable and can reduce the gap, without deriving any inequality.
-
other
[Section V-A (Theoretical Analysis of H\"older Divergence), Theorem 2 proof; verbatim duplicate as Theorem 3]
"To show that the ELBO with the PHD is tighter than the ELBO with the KL divergence, we need to show that: ELBO H ≥ ELBOKL. Since the PHD is more flexible and tunable through the parameters α, β, γ, it can better fit the true posterior distribution and reduce the gap between the variational distribution and the true posterior."
The theorem to be proved is exactly ELBO_H ≥ ELBO_KL. The proof's only substantive step is to assert that PHD can 'better fit the true posterior distribution and reduce the gap'—the very statement at issue. No inequality relating D_PHD to D_KL, and no comparison of the full variational objectives in Eqs. (19)-(20) versus Eq. (3), is derived. 'Tunable through the parameters α, β, γ' is a property of the PHD definition, not a proof that any particular (α,β,γ) yields a tighter bound. Theorem 3 repeats the same non-proof verbatim, so the central 'theoretical guarantee' is the conclusion restated as its own justification.
-
other
[Section V-A, Theorem 4 proof]
"By adjusting the parameters α, β, and γ, we can ensure that the divergence measures the entire distribution rather than collapsing to a single mode."
The theorem claims that PHD avoids mode collapse because it is more flexible than KL. The proof's operative sentence asserts that adjusting α, β, γ 'ensure[s]' the desired behavior, which is exactly the theorem's conclusion. No construction of parameters, no divergence inequality, and no comparison of the variational objectives is supplied. This is a restatement of the desired property rather than a derivation from the PHD formula, reinforcing the same circular step found in Theorem 2.
full rationale
The paper's computational machinery is not circular: the closed-form PHD for Dirichlet distributions in Theorem 1 is imported from the independent source [9] via Lemmas 1-2, and the gamma-function expansions in Eqs. (14)-(16) are standard identities rather than redefinitions of the target result. The empirical comparison also does not exhibit constructional circularity; per-dataset grid selection of α and γ (Tables III, V, IX) is a post-hoc selection concern, not a fitted parameter being renamed as a prediction. The circularity is concentrated in the advertised theoretical guarantee. Theorem 2 states that PHD gives a tighter ELBO than KL, and its proof, after writing the two objectives, argues only that PHD's tunable α, β, γ 'can better fit the true posterior distribution and reduce the gap'—the very inequality the theorem must establish. Theorem 3 duplicates this argument verbatim, and Theorem 4 justifies mode-collapse avoidance with the same tunability assertion. These are not derivations; the central theoretical claim reduces to a restatement of itself. Because no load-bearing self-citation is present and the PHD has independent support from [9], the score is 6 rather than higher, but the paper's central theoretical result is circular in its proof structure.
Assumptions & free parameters
free parameters (4)
- alpha (Hölder exponent) =
1.1 to 2.5 per dataset
- gamma (Hölder gamma exponent) =
0.5 to 2.0 per dataset
- lambda_t (regularization weight) =
not reported
- Kalman filter noise covariances =
not reported
assumptions (6)
- standard math Lemma 1: PHD has closed form for conic or affine exponential families (from Nielsen et al. [9]).
- standard math Dirichlet distribution is an exponential family with natural parameter alpha and log-normalizer F(theta) = sum log Gamma(theta_k+1) - log Gamma(sum(theta_k+1)).
- ad hoc to paper Replacing KL with PHD in the variational objective preserves a valid evidence lower bound and yields ELBO_H >= ELBO_KL.
- ad hoc to paper Gamma-function expansions in Eqs. (14)-(16) are valid as written.
- domain assumption Simplified Dempster-Shafer fusion rules assume independence of views.
- ad hoc to paper Pseudo-views formed by concatenating outputs of two backbones are valid views for clustering.
Cite this review
Pith. "Pith review of Uncertainty Quantification for Incomplete Multi-View Data Using Divergence Measures." pith.science (2026). https://pith.science/paper/6NTHOAUE
@misc{pith2026250709980,
author = {Pith},
title = {Pith review of: Uncertainty Quantification for Incomplete Multi-View Data Using Divergence Measures},
year = {2026},
howpublished = {\url{https://pith.science/paper/6NTHOAUE}},
note = {Machine review of arXiv:2507.09980}
}
read the original abstract
Existing multi-view classification and clustering methods typically improve task accuracy by leveraging and fusing information from different views. However, ensuring the reliability of multi-view integration and final decisions is crucial, particularly when dealing with noisy or corrupted data. Current methods often rely on Kullback-Leibler (KL) divergence to estimate uncertainty of network predictions, ignoring domain gaps between different modalities. To address this issue, KPHD-Net, based on H\"older divergence, is proposed for multi-view classification and clustering tasks. Generally, our KPHD-Net employs a variational Dirichlet distribution to represent class probability distributions, models evidences from different views, and then integrates it with Dempster-Shafer evidence theory (DST) to improve uncertainty estimation effects. Our theoretical analysis demonstrates that Proper H\"older divergence offers a more effective measure of distribution discrepancies, ensuring enhanced performance in multi-view learning. Moreover, Dempster-Shafer evidence theory, recognized for its superior performance in multi-view fusion tasks, is introduced and combined with the Kalman filter to provide future state estimations. This integration further enhances the reliability of the final fusion results. Extensive experiments show that the proposed KPHD-Net outperforms the current state-of-the-art methods in both classification and clustering tasks regarding accuracy, robustness, and reliability, with theoretical guarantees.
Figures
Reference graph
Works this paper leans on
-
[1]
Trusted multi-view classi- fication with dynamic evidential fusion,
Z. Han, C. Zhang, H. Fu, and J. T. Zhou, “Trusted multi-view classi- fication with dynamic evidential fusion,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 2551–2566, 2022
work page 2022
-
[2]
Incomplete contrastive multi-view clustering with high-confidence guiding,
G. Chao, Y . Jiang, and D. Chu, “Incomplete contrastive multi-view clustering with high-confidence guiding,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 10, 2024, pp. 11 221– 11 229
work page 2024
-
[3]
Exploring and Exploiting Uncertainty for Incomplete Multi-View Classification,
M. Xie, Z. Han, C. Zhang, Y . Bai, and Q. Hu, “Exploring and Exploiting Uncertainty for Incomplete Multi-View Classification,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023 . IEEE, 2023, pp. 19 873– 19 882
work page 2023
-
[4]
Fast continual multi-view clustering with incomplete views,
X. Wan, B. Xiao, X. Liu, J. Liu, W. Liang, and E. Zhu, “Fast continual multi-view clustering with incomplete views,” IEEE Transactions on Image Processing, vol. 33, pp. 2995–3008, 2024
work page 2024
-
[5]
Self-supervised information bottleneck for deep multi-view subspace clustering,
S. Wang, C. Li, Y . Li, Y . Yuan, and G. Wang, “Self-supervised information bottleneck for deep multi-view subspace clustering,” IEEE Transactions on Image Processing , vol. 32, pp. 1555–1567, 2023. IEEE TRANSACTIONS ON IMAGE PROCESSING 15
work page 2023
-
[6]
Posterior network: Uncertainty estimation without ood samples via density-based pseudo- counts,
B. Charpentier, D. Z ¨ugner, and S. G ¨unnemann, “Posterior network: Uncertainty estimation without ood samples via density-based pseudo- counts,” Advances in Neural Information Processing Systems , vol. 33, pp. 1356–1367, 2020
work page 2020
-
[8]
On the Dempster-Shafer framework and new combination rules,
R. R. Yager, “On the Dempster-Shafer framework and new combination rules,” Information Sciences, vol. 41, no. 2, pp. 93–137, 1987
work page 1987
-
[9]
On H ¨older projective divergences,
F. Nielsen, K. Sun, and S. Marchand-Maillet, “On H ¨older projective divergences,” Entropy, vol. 19, no. 3, p. 122, 2017
work page 2017
Show all 77 references
-
[10]
Bayesian filtering: From Kalman filters to particle filters, and beyond,
Z. Chen et al., “Bayesian filtering: From Kalman filters to particle filters, and beyond,” Statistics, vol. 182, no. 1, pp. 1–69, 2003
2003
-
[11]
I-divergence geometry of probability distributions and min- imization problems,
I. Csisz ´ar, “I-divergence geometry of probability distributions and min- imization problems,” The Annals of Probability , pp. 146–158, 1975
1975
-
[12]
Self-supervised geometric features discovery via interpretable attention for vehicle re-identification and beyond,
M. Li, X. Huang, and Z. Zhang, “Self-supervised geometric features discovery via interpretable attention for vehicle re-identification and beyond,” in ICCV, 2021
2021
-
[13]
Exploiting multi- view part-wise correlation via an efficient transformer for vehicle re- identification,
M. Li, J. Liu, C. Zheng, X. Huang, and Z. Zhang, “Exploiting multi- view part-wise correlation via an efficient transformer for vehicle re- identification,” TOM, 2021
2021
-
[14]
Synthetic-to-real self-supervised robust depth estimation via learning with motion and structure priors,
W. Yan, M. Li, H. Li, S. Shao, and R. T. Tan, “Synthetic-to-real self-supervised robust depth estimation via learning with motion and structure priors,” in Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR) , June 2025, pp. 21 880–21 890
2025
-
[15]
Eventgpt: Event stream understanding with multimodal large language models,
S. Liu, J. Li, G. Zhao, Y . Zhang, X. Meng, F. R. Yu, X. Ji, and M. Li, “Eventgpt: Event stream understanding with multimodal large language models,” in Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), June 2025, pp. 29 139–29 149
2025
-
[16]
Favchat: Unlocking fine-grained facial video understanding with multimodal large language models,
F. Zhao, M. Li, L. Xu, W. Jiang, J. Gao, and D. Yan, “Favchat: Unlocking fine-grained facial video understanding with multimodal large language models,” arXiv preprint arXiv:2503.09158 , 2025
2025 arXiv
-
[17]
Unif2ace: Fine-grained face understanding and generation with unified multimodal models,
J. Li, X. Qiu, L. Xu, L. Guo, D. Qu, T. Long, C. Fan, and M. Li, “Unif2ace: Fine-grained face understanding and generation with unified multimodal models,” arXiv preprint arXiv:2503.08120 , 2025
2025
-
[18]
Dr-fer: Discriminative and robust representation learning for facial expression recognition,
M. Li, H. Fu, S. He, H. Fan, J. Liu, J. Keppo, and M. Z. Shou, “Dr-fer: Discriminative and robust representation learning for facial expression recognition,” IEEE Transactions on Multimedia, vol. 26, pp. 6297–6309, 2023
2023
-
[19]
Stprivacy: Spatio-temporal privacy-preserving action recognition,
M. Li, X. Xu, H. Fan, P. Zhou, J. Liu, J.-W. Liu, J. Li, J. Keppo, M. Z. Shou, and S. Yan, “Stprivacy: Spatio-temporal privacy-preserving action recognition,” in ICCV, 2023
2023
-
[20]
Instant3d: instant text-to-3d generation,
M. Li, P. Zhou, J.-W. Liu, J. Keppo, M. Lin, S. Yan, and X. Xu, “Instant3d: instant text-to-3d generation,” IJCV, 2024
2024
-
[21]
Realera: Semantic-level concept erasure via neighbor-concept mining,
Y . Liu, J. An, W. Zhang, M. Li, D. Wu, J. Gu, Z. Lin, and W. Wang, “Realera: Semantic-level concept erasure via neighbor-concept mining,” arXiv preprint arXiv:2410.09140 , 2024
2024 arXiv
-
[22]
Vistorybench: Comprehensive bench- mark suite for story visualization,
C. Zhuang, A. Huang, W. Cheng, J. Wu, Y . Hu, J. Liao, Z. Huang, H. Wang, X. Liao, W. Cai et al., “Vistorybench: Comprehensive bench- mark suite for story visualization,” arXiv preprint arXiv:2505.24862 , 2025
2025
-
[23]
Reliable conflictive multi-view learning,
C. Xu, J. Si, Z. Guan, W. Zhao, Y . Wu, and X. Gao, “Reliable conflictive multi-view learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 14, 2024, pp. 16 129–16 137
2024
-
[24]
One-Step Multi-View Clustering With Diverse Representation,
X. Wan, J. Liu, X. Gan, X. Liu, S. Wang, Y . Wen, T. Wan, and E. Zhu, “One-Step Multi-View Clustering With Diverse Representation,” IEEE Transactions on Neural Networks and Learning Systems , vol. 36, no. 3, pp. 5774–5786, 2025
2025
-
[25]
Multi-view multi-label canonical correlation analysis for cross-modal matching and retrieval,
R. Sanghavi and Y . Verma, “Multi-view multi-label canonical correlation analysis for cross-modal matching and retrieval,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 4701–4710
2022
-
[26]
A multi-view co-training network for semi-supervised medical image-based prognostic prediction,
H. Li, S. Wang, B. Liu, M. Fang, R. Cao, B. He, S. Liu, C. Hu, D. Dong, X. Wang et al., “A multi-view co-training network for semi-supervised medical image-based prognostic prediction,” Neural Networks, vol. 164, pp. 455–463, 2023
2023
-
[27]
Visual tracking with spatio-temporal Dempster–Shafer information fusion,
X. Li, A. Dick, C. Shen, Z. Zhang, A. van den Hengel, and H. Wang, “Visual tracking with spatio-temporal Dempster–Shafer information fusion,” IEEE Transactions on Image Processing , vol. 22, no. 8, pp. 3028–3040, 2013
2013
-
[28]
Variational Quantum Linear Solver-based Combination Rules in Dempster–Shafer Theory,
H. Luo, Q. Zhou, Z. Li, and Y . Deng, “Variational Quantum Linear Solver-based Combination Rules in Dempster–Shafer Theory,” Informa- tion Fusion, vol. 102, p. 102070, 2024
2024
-
[29]
Scale-invariant divergences for density functions,
T. Kanamori, “Scale-invariant divergences for density functions,” En- tropy, vol. 16, no. 5, pp. 2611–2628, 2014
2014
-
[30]
The Cauchy– Schwarz divergence for Poisson point processes,
H. G. Hoang, B.-N. V o, B.-T. V o, and R. Mahler, “The Cauchy– Schwarz divergence for Poisson point processes,” IEEE Transactions on Information Theory , vol. 61, no. 8, pp. 4475–4485, 2015
2015
-
[31]
K-means clustering with H ¨older divergences,
F. Nielsen, K. Sun, and S. Marchand-Maillet, “K-means clustering with H ¨older divergences,” in Geometric Science of Information: Third International Conference, GSI 2017, Paris, France, November 7-9, 2017, Proceedings 3. Springer, 2017, pp. 856–863
2017
-
[32]
Classification for Polsar image based on h ¨older divergences,
T. Pan, D. Peng, X. Yang, P. Huang, and W. Yang, “Classification for Polsar image based on h ¨older divergences,” The Journal of Engineering, vol. 2019, no. 21, pp. 7593–7596, 2019
2019
-
[33]
Inverse kalman filtering problems for discrete-time systems,
Y . Li, B. Wahlberg, X. Hu, and L. Xie, “Inverse kalman filtering problems for discrete-time systems,” Automatica, vol. 163, p. 111560, 2024
2024
-
[34]
Adaptive kalman filtering for histogram-based appearance learn- ing in infrared imagery,
V . Venkataraman, G. Fan, J. P. Havlicek, X. Fan, Y . Zhai, and M. B. Yeary, “Adaptive kalman filtering for histogram-based appearance learn- ing in infrared imagery,” IEEE Transactions on Image Processing , vol. 21, no. 11, pp. 4622–4635, 2012
2012
-
[35]
Efficient and fast real-world noisy image denoising by combining pyramid neural network and two-pathway unscented Kalman filter,
R. Ma, H. Hu, S. Xing, and Z. Li, “Efficient and fast real-world noisy image denoising by combining pyramid neural network and two-pathway unscented Kalman filter,” IEEE Transactions on Image Processing , vol. 29, pp. 3927–3940, 2020
2020
-
[36]
Kalman filter for spatial- temporal regularized correlation filters,
S. Feng, K. Hu, E. Fan, L. Zhao, and C. Wu, “Kalman filter for spatial- temporal regularized correlation filters,” IEEE Transactions on Image Processing, vol. 30, pp. 3263–3278, 2021
2021
-
[37]
The statistical analysis of compositional data,
J. Aitchison, “The statistical analysis of compositional data,” Journal of the Royal Statistical Society: Series B (Methodological) , vol. 44, no. 2, pp. 139–160, 1982
1982
-
[38]
Dirichlet and related distribu- tions: Theory, methods and applications,
K. W. Ng, G.-L. Tian, and M.-L. Tang, “Dirichlet and related distribu- tions: Theory, methods and applications,” 2011
2011
-
[39]
C. M. Bishop and N. M. Nasrabadi, Pattern recognition and machine learning. Springer, 2006, vol. 4, no. 4
2006
-
[40]
Clustering documents with an exponential-family approxima- tion of the Dirichlet compound multinomial distribution,
C. Elkan, “Clustering documents with an exponential-family approxima- tion of the Dirichlet compound multinomial distribution,” in Proceedings of the 23rd International Conference on Machine Learning , 2006, pp. 289–296
2006
-
[41]
Jsang, Subjective Logic: A formalism for reasoning under uncertainty
A. Jsang, Subjective Logic: A formalism for reasoning under uncertainty. Springer Publishing Company, Incorporated, 2018
2018
-
[42]
Belief functions and parametric models,
G. Shafer, “Belief functions and parametric models,” Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 44, no. 3, pp. 322–339, 1982
1982
-
[43]
A Dempster-Shafer Approach to Trustworthy AI With Application to Fetal Brain MRI Segmentation,
L. Fidon, M. Aertsen, F. Kofler, A. Bink, A. L. David, T. Deprest, D. Emam, F. Guffens, A. Jakab, G. Kasprian, P. Kienast, A. Melbourne, B. Menze, N. Mufti, I. Pogledic, D. Prayer, M. Stuempflen, E. Van Els- lander, S. Ourselin, J. Deprest, and T. Vercauteren, “A Dempster-Shaf...
2024
-
[44]
Multimodal Dy- namics: Dynamical Fusion for Trustworthy Multimodal Classification,
Z. Han, F. Yang, J. Huang, C. Zhang, and J. Yao, “Multimodal Dy- namics: Dynamical Fusion for Trustworthy Multimodal Classification,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022 . IEEE, 2022, pp. 20 675–20 685
2022
-
[45]
Cauchy–Schwarz Regularized Autoencoder,
L. Tran, M. Pantic, and M. P. Deisenroth, “Cauchy–Schwarz Regularized Autoencoder,” Journal of Machine Learning Research, vol. 23, no. 115, pp. 1–37, 2022
2022
-
[46]
Sun RGB-D: A RGB-D scene understanding benchmark suite,
S. Song, S. P. Lichtenberg, and J. Xiao, “Sun RGB-D: A RGB-D scene understanding benchmark suite,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 567–576
2015
-
[47]
Indoor segmentation and support inference from rgbd images,
N. Silberman, D. Hoiem, P. Kohli, and R. Fergus, “Indoor segmentation and support inference from rgbd images,” in Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part V 12 . Springer, 2012, pp. 746– 760
2012
-
[48]
Semantic understanding of scenes through the ade20k dataset,
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ade20k dataset,” International Journal of Computer Vision , vol. 127, pp. 302– 321, 2019
2019
-
[49]
Scannet: Richly-annotated 3D reconstructions of indoor scenes,
A. Dai, A. X. Chang, M. Savva, M. Halber, T. Funkhouser, and M. Nießner, “Scannet: Richly-annotated 3D reconstructions of indoor scenes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5828–5839
2017
-
[50]
Parameter-free auto-weighted multiple graph learning: A framework for multiview clustering and semi-supervised classification
F. Nie, J. Li, X. Li et al., “Parameter-free auto-weighted multiple graph learning: A framework for multiview clustering and semi-supervised classification.” in IJCAI, vol. 9, 2016, pp. 1881–1887
2016
-
[51]
A fast iterative shrinkage-thresholding algo- rithm for linear inverse problems,
A. Beck and M. Teboulle, “A fast iterative shrinkage-thresholding algo- rithm for linear inverse problems,” SIAM Journal on Imaging Sciences , vol. 2, no. 1, pp. 183–202, 2009
2009
-
[52]
Unsupervised video matting via sparse and low-rank representation,
D. Zou, X. Chen, G. Cao, and X. Wang, “Unsupervised video matting via sparse and low-rank representation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 6, pp. 1501–1514, 2019. IEEE TRANSACTIONS ON IMAGE PROCESSING 16
2019
-
[54]
DasGupta, Probability for statistics and machine learning: fundamentals and advanced topics
A. DasGupta, Probability for statistics and machine learning: fundamentals and advanced topics. Springer, 2011. [Online]. Available: https://api.semanticscholar.org/CorpusID:124734892
2011
-
[55]
Learning deep sparse regu- larizers with applications to multi-view clustering and semi-supervised classification,
S. Wang, Z. Chen, S. Du, and Z. Lin, “Learning deep sparse regu- larizers with applications to multi-view clustering and semi-supervised classification,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 9, pp. 5042–5055, 2022
2022
-
[56]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770–778
2016
-
[57]
ImageNet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2009, pp. 248–255
2009
-
[58]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international con- ference, Munich, Germany, October 5-9, 2015, proceedings, part III 18 ...
2015
-
[59]
Densely connected convolutional networks,
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition, 2017, pp. 4700–4708
2017
-
[60]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in 9th International Conference on Lear...
2021
-
[61]
Hungry Hungry Hippos: Towards Language Modeling with State Space Models,
D. Y . Fu, T. Dao, K. K. Saab, A. W. Thomas, A. Rudra, and C. R ´e, “Hungry Hungry Hippos: Towards Language Modeling with State Space Models,” in The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023
2023
-
[62]
Adam: A Method for Stochastic Optimiza- tion,
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimiza- tion,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015
2015
-
[63]
Pytorch: An imperative style, high-performance deep learning library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
-
[64]
Translate-to- recognize networks for RGB-D scene recognition,
D. Du, L. Wang, H. Wang, K. Zhao, and G. Wu, “Translate-to- recognize networks for RGB-D scene recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 11 836–11 845
2019
-
[65]
Image rep- resentations with spatial object-to-object relations for RGB-D scene recognition,
X. Song, S. Jiang, B. Wang, C. Chen, and G. Chen, “Image rep- resentations with spatial object-to-object relations for RGB-D scene recognition,” IEEE Transactions on Image Processing, vol. 29, pp. 525– 537, 2019
2019
-
[66]
Centroid Based Concept Learning for RGB-D Indoor Scene Classification,
A. Ayub and A. R. Wagner, “Centroid Based Concept Learning for RGB-D Indoor Scene Classification,” in 31st British Machine Vision Conference 2020, BMVC 2020, Virtual Event, UK, September 7-10,
2020
-
[67]
Cross-modal pyramid translation for RGB-D scene recognition,
D. Du, L. Wang, Z. Li, and G. Wu, “Cross-modal pyramid translation for RGB-D scene recognition,” International Journal of Computer Vision , vol. 129, no. 8, pp. 2309–2327, 2021
2021
-
[68]
When CNNs meet random RNNs: Towards multi-level analysis for RGB-D object and scene recognition,
A. Caglayan, N. Imamoglu, A. B. Can, and R. Nakamura, “When CNNs meet random RNNs: Towards multi-level analysis for RGB-D object and scene recognition,” Computer Vision and Image Understanding, vol. 217, p. 103373, 2022
2022
-
[69]
FGMNet: Feature grouping mechanism network for RGB-D indoor scene semantic seg- mentation,
Y . Zhang, W. Zhou, L. Ye, L. Yu, and T. Luo, “FGMNet: Feature grouping mechanism network for RGB-D indoor scene semantic seg- mentation,” Digital Signal Processing , vol. 149, p. 104480, 2024
2024
-
[70]
Feature contrast difference and enhanced network for rgb-d indoor scene classification in internet of things,
W. Zhou, B. Jian, and Y . Liu, “Feature contrast difference and enhanced network for rgb-d indoor scene classification in internet of things,” IEEE Internet of Things Journal , vol. 12, no. 11, pp. 17 610–17 621, 2025
2025
-
[71]
An efficient k-means clustering algorithm: Analysis and implementation,
T. Kanungo, D. M. Mount, N. S. Netanyahu, C. D. Piatko, R. Silverman, and A. Y . Wu, “An efficient k-means clustering algorithm: Analysis and implementation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 24, no. 7, pp. 881–892, 2002
2002
-
[72]
Self-weighted multiview clustering with multiple graphs
F. Nie, J. Li, X. Li et al. , “Self-weighted multiview clustering with multiple graphs.” in IJCAI, 2017, pp. 2564–2570
2017
-
[73]
Multi-view clustering and semi-supervised classification with adaptive neighbours,
F. Nie, G. Cai, and X. Li, “Multi-view clustering and semi-supervised classification with adaptive neighbours,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 31, no. 1, 2017
2017
-
[74]
Multiview consensus graph clustering,
K. Zhan, F. Nie, J. Wang, and Y . Yang, “Multiview consensus graph clustering,” IEEE Transactions on Image Processing , vol. 28, no. 3, pp. 1261–1270, 2018
2018
-
[75]
Binary multi- view clustering,
Z. Zhang, L. Liu, F. Shen, H. T. Shen, and L. Shao, “Binary multi- view clustering,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 7, pp. 1774–1782, 2018
2018
-
[76]
Multi-view subspace clustering with intactness-aware similarity,
X. Wang, Z. Lei, X. Guo, C. Zhang, H. Shi, and S. Z. Li, “Multi-view subspace clustering with intactness-aware similarity,” Pattern Recogni- tion, vol. 88, pp. 50–63, 2019
2019
-
[77]
Multi-view clustering based on view-attention driven,
Z. Ma, J. Yu, L. Wang, H. Chen, Y . Zhao, X. He, Y . Wang, and Y . Song, “Multi-view clustering based on view-attention driven,” International Journal of Machine Learning and Cybernetics , vol. 14, no. 8, pp. 2621– 2631, 2023
2023
-
[78]
One for all: A novel Dual-space Co-training baseline for Large-scale Multi-View Clustering,
Z. Kong, Z. Fu, D. Chang, Y . Wang, and Y . Zhao, “One for all: A novel Dual-space Co-training baseline for Large-scale Multi-View Clustering,” arXiv preprint arXiv:2401.15691 , 2024
2024 arXiv
-
[79]
Deep Incomplete Multi-View Learning Network with Insufficient Label Information,
Z. Jiang, T. Luo, and X. Liang, “Deep Incomplete Multi-View Learning Network with Insufficient Label Information,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 11, 2024, pp. 12 919–12 927
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.