REVIEW 5 major objections 5 minor 62 references
SC-GIR: Goal-oriented Semantic Communication via Invariant Representation Learning
T0 review · 5 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read SC-GIR claims a self-supervised, covariance-based encoder can extract task-essential image features, transmit them over noisy channels, and still classify above 85% accuracy at high compression.
desk verdict A sensible but incremental application of Barlow Twins to semantic communication; the benchmark is broad and the SNR curves are suggestive, but the paper's own numbers are inconsistent and the channel codec protocol is too underspecified to support the headline claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the cross-correlation loss on standardized embeddings: L_cross-corr = sum_i (1 - C_ii)^2 + λ sum_{i≠j} (C_ij)^2, where C is the empirical cross-correlation matrix between the outputs of two parallel encoders fed with two augmented views of the same image. The diagonal terms push the representation to be invariant to the augmentations, and the off-diagonal terms push different feature dimensions to be decorrelated, which the paper interprets as redundancy reduction aligned with the information bottleneck principle. This loss is computed during training only; at inference, a single non-augmented image is encoded once and transmitted, so the framework adds no infer
What would settle it
At k/n = 0.1 and SNR = 5 dB, measure the actual transmitted energy per image and the number of channel uses for SC-GIR, DeepJSCC, SemCC, SemRE, and BPG+LDPC, then rerun the classification comparison with exactly matched power and bandwidth. If SC-GIR's 85% AWGN and 80% Rayleigh accuracy advantage disappears under matched resource budgets, the headline outperformance claim is an artifact of unequal comparison rather than a property of the learned representation.
Extended reading notes
Core claim
The central claim is that a goal-oriented semantic communication system can be built from an invariant representation learned purely by self-supervision: two distorted views of the same image are fed through the same encoder, and a covariance-based cross-correlation loss makes the embedding stable across views while decorrelating its dimensions. This yields a compressed latent that survives Rayleigh fading and AWGN well enough for downstream classification, without requiring joint training of transmitter and receiver or labeled data. The paper further claims that the same representation transfers to semantic segmentation and domain generalization, reporting 63.5 mean IoU on Cityscapes and 76
Load-bearing premise
The reported accuracy gains assume that every method transmits the same number of channel symbols under the same power and bandwidth budget at each compression ratio, but the paper never specifies the channel encoder/decoder architectures or power normalization used for the baselines.
Editorial extensions
If this is right
- If SC-GIR is correct, semantic communication no longer requires labeled training data or joint transmitter-receiver training, removing a major barrier to deployment in dynamic IoT environments.
- The reported accuracy at k/n = 0.1 suggests that a tenfold bandwidth reduction is possible for classification-oriented image transmission while keeping task accuracy above 80-85% on common benchmarks.
- Because the same representation supports classification, segmentation, and domain-generalization tasks, the framework points toward one encoder serving many downstream goals at the receiver rather than a dedicated codec per task.
- The training-time-only augmentation scheme means the learned invariance comes at no extra inference cost, making the approach compatible with resource-constrained edge devices.
- If the domain-generalization results hold, SC-GIR could reduce or eliminate the need to retrain the semantic encoder when a deployment environment's visual style changes.
Reading between the lines
- If the representation is truly task-agnostic, one could test it by attaching multiple downstream heads (classifier, segmenter, detector) to the same transmitted latent and measuring whether all benefit without per-task retraining of the encoder; the paper evaluates classification and segmentation separately but not simultaneously.
- The cross-correlation loss depends on batch statistics for standardization, so a natural stress test is to evaluate SC-GIR with very small or non-i.i.d. batches, which are common in real-time edge inference; the paper's experiments use larger batches and do not address this failure mode.
- The 'invariant' claim could be probed by causally shifting spurious correlations in the source data, e.g., changing background or lighting while keeping the label fixed; PACS is a useful domain-shift proxy but not a causal intervention, so the invariance claim remains partially open.
- The paper suggests the framework extends to text and audio; since the covariance loss is modality-agnostic, a concrete next step would be a multimodal benchmark with the same Rayleigh channel model, testing whether the same redundancy-reduction principle transfers across signal types.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SC-GIR, a self-supervised semantic communication framework for image transmission. A ResNet-34 encoder with a multi-layer projection head is trained with a Barlow-Twins-style cross-correlation loss on two augmented views, producing a compressed latent that is transmitted over AWGN/Rayleigh channels and used for downstream classification. Experiments on CIFAR-10/100, MNIST/FMNIST, STL-10, Flower-17, Cityscapes, and PACS are reported, and the abstract claims nearly 10% improvement over baselines and over 85% classification accuracy at low SNR/compression. The paper also includes segmentation and domain-generalization experiments as evidence of task-agnostic representation quality.
Significance. If the resource-controlled comparison and the reported numbers were correct, the paper would be a useful demonstration that a self-supervised covariance-based encoder can serve as a task-agnostic semantic source for wireless image classification, reducing reliance on labeled data and joint end-to-end training. The claim of robust performance at low compression ratios and low SNR is practically significant for IoT/edge communications. However, the manuscript currently contains several load-bearing inconsistencies: the headline average accuracy in Table V is not the mean of the reported rows, Eq. (3) is not a valid information-theoretic identity, and the channel-symbol/power budgets against baselines are not specified. These issues must be resolved before the central claims can be accepted.
major comments (5)
- [Abstract and Table V] The abstract claims SC-GIR outperforms baselines by 'nearly 10%', but Table V's reported rows contradict this. The SC-GIR row average is (87.2+98.0+99.3+85.5+86.5+73.1)/6 ≈ 88.3, not 66.6; DeepJSCC's row average is ≈ 73.7, not 66.7. The 'Average' column is therefore not the mean of the per-dataset entries. The Section V-B discussion ('average accuracy of 66.6% ... nearly matching DeepJSCC and DeepSC, both at 66.7%') would imply no gain, not a 10% gain. Please recompute the table, restate the headline claim, and reconcile the abstract with the corrected numbers.
- [Eq. (3), Section III-B] The decomposition I(X;S) = I(X;S|Y) + I(S;Y) is not an identity in general. The chain rule gives I(X;S) = I(X;S|Y) + I(S;Y) only under additional assumptions (e.g., some form of conditional independence/dependence structure that is not stated). This equation is the formal justification for the redundant/task-related split and for the subsequent causal model. Please replace it with a correct decomposition or explicitly state and justify the assumptions under which it holds; otherwise the theoretical motivation for the loss is unsupported.
- [Section III-D and Section V-C] The comparison with baselines is not resource-controlled as reported. The paper never specifies the channel encoder/decoder architecture (Section III-D is qualitative), how the 2048-dimensional semantic latent is mapped to the stated k/n ratios and transmitted byte sizes (307–1,843 bytes), or the transmit power normalization used to set SNR in Eq. (1). Different methods may therefore transmit different numbers of channel symbols or at different average power for the same nominal k/n. This directly affects the validity of Figs. 5–6 and the 'outperforms baselines' claims. Please specify the full transmission chain, including codec, symbol mapping, and power normalization, and verify that all methods use equal channel resources.
- [Section IV-B, Eqs. (10) and (13)] The paper states that the IB objective in Eq. (13) is 'reformulated' into the cross-correlation loss in Eq. (10) 'through simplifications and approximations [40]'. This is not demonstrated and is questionable: Eq. (10) is the Barlow Twins loss, a heuristic covariance-based objective, and no derivation connecting the IB Lagrangian to the diagonal/off-diagonal cross-correlation terms is given. Please either provide a concrete derivation or reframe Eq. (10) as an empirically motivated loss, without claiming it is derived from IB.
- [Section V-F, Tables VIII-IX] The generalization claims in Section V-F are not supported by sufficient experimental detail. For the Cityscapes segmentation experiment, the paper does not state the decoder architecture, training protocol, resolution, or how SC-GIR's semantic encoder is integrated; for PACS, the fine-tuning/evaluation protocol is not given. Without these details, the strong mIoU and domain-generalization numbers cannot be assessed. Please add the missing protocol information or temper the generalization claims accordingly.
minor comments (5)
- [Algorithm 1] The algorithm does not match Eq. (10): the λ weighting on the off-diagonal term is missing, and the definitions of Lon and Loff (e.g., line 8) are unclear—the notation 'C − fdiag(C) + 1' is not a standard way to target off-diagonal entries. Please align the pseudocode with the equation.
- [Section IV-B, paragraph after Eq. (10)] The text says 'The first component of the cross-correlation focuses on the off-diagonal loss', but Eq. (10)'s first term is the diagonal term (1−Cii)². This appears to be a wording error.
- [Fig. 5 caption] The Rayleigh subplot caption says 'BPG 12 rate LDPC' while the text refers to 'BPG 3/4 rate LDPC'. Please correct the inconsistency.
- [Section V-A, Metrics] The metric description says 'High cosine similarity between the original data and encoded representations', but Fig. 8 shows cosine similarity between latent representations of the two augmented views, not between original data and encoding. Please clarify.
- [References] Reference [59] appears to duplicate reference [26]; both cite the same 'Contrastive learning-based semantic communications' paper. Please consolidate.
Circularity Check
Main benchmark results are external and non-circular, but the 'Encoder Evaluation' plots report the training objective itself as evidence.
-
self definitional
[Section V-D 'Encoder Evaluation', Fig. 8 and Fig. 9; Eq. (10)]
"Fig. 8 depicts the cosine similarity between latent representations derived from the proposed semantic encoder during training for different datasets. The noticeable increase in cosine similarity indicates the model’s increasing resilience to augmentations applied to the two views."
The SC-GIR loss in Eq. (10) is Lcross-corr = Σ_i (1 − C_ii)^2 + λ Σ_i Σ_{i≠j} (C_ij)^2, so minimizing it directly drives the diagonal cross-correlation (i.e., the cosine similarity between the two augmented-view embeddings) toward 1 and off-diagonal entries toward 0. Therefore the increasing cosine similarity in Fig. 8 and the diagonalization of the cross-correlation matrices in Fig. 9 are exactly the training objective being satisfied, not independent evidence of semantic quality or noise resilience. The evaluation metric reduces by construction to the quantity being optimized in the loss.
full rationale
The central empirical claims—classification accuracy against SemCC, SemRE, DeepJSCC, DeepSC, and BPG+LDPC—are benchmarked against external methods, so those results are not circular. The information-bottleneck motivation (Eq. 2) is not derived rigorously; the cross-correlation objective (Eq. 10) is explicitly imported from Barlow Twins [40] via 'simplifications and approximations,' which is an attribution/novelty issue rather than a circular derivation. Self-citations such as [1] and [4] are background and not load-bearing. The one self-referential element is the Section V-D 'Encoder Evaluation': Figs. 8 and 9 plot cosine similarity and cross-correlation, which Eq. (10) is explicitly optimized to drive to 1 and 0, so those plots show the loss converging rather than an independent success measure. This does not affect the main benchmark comparisons, so overall circularity is low. Separate correctness risks, not counted as circularity, include the inconsistent Average column in Table V and the unsupported attribution of robustness to 'generative image restoration' in Section V-C.
Assumptions & free parameters
free parameters (3)
- lambda (off-diagonal loss weight) =
5e-4
- beta in IB objective Eq. (2) =
not specified
- alpha in IB objective Eq. (13) =
not specified
assumptions (5)
- ad hoc to paper I(X;S) = I(X;S|Y) + I(S;Y) (Eq. 3)
- ad hoc to paper The cross-correlation loss Eq. (10) is a simplification of the IB objective Eq. (13)
- domain assumption Causal and non-causal parts are disjoint and C is the sole parent of label L (Definition 1)
- domain assumption G-invariance P(S|X)=P(S|T·X) implies task-relevant minimal sufficiency
- domain assumption The Rayleigh fading channel coefficient h is available or learnable at the receiver without further specification
Cite this review
Pith. "Pith review of SC-GIR: Goal-oriented Semantic Communication via Invariant Representation Learning." pith.science (2026). https://pith.science/paper/OHM2IWNA
@misc{pith2026250901119,
author = {Pith},
title = {Pith review of: SC-GIR: Goal-oriented Semantic Communication via Invariant Representation Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/OHM2IWNA}},
note = {Machine review of arXiv:2509.01119}
}
read the original abstract
Goal-oriented semantic communication (SC) aims to revolutionize communication systems by transmitting only task-essential information. However, current approaches face challenges such as joint training at transceivers, leading to redundant data exchange and reliance on labeled datasets, which limits their task-agnostic utility. To address these challenges, we propose a novel framework called Goal-oriented Invariant Representation-based SC (SC-GIR) for image transmission. Our framework leverages self-supervised learning to extract an invariant representation that encapsulates crucial information from the source data, independent of the specific downstream task. This compressed representation facilitates efficient communication while retaining key features for successful downstream task execution. Focusing on machine-to-machine tasks, we utilize covariance-based contrastive learning techniques to obtain a latent representation that is both meaningful and semantically dense. To evaluate the effectiveness of the proposed scheme on downstream tasks, we apply it to various image datasets for lossy compression. The compressed representations are then used in a goal-oriented AI task. Extensive experiments on several datasets demonstrate that SC-GIR outperforms baseline schemes by nearly 10%,, and achieves over 85% classification accuracy for compressed data under different SNR conditions. These results underscore the effectiveness of the proposed framework in learning compact and informative latent representations.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[40]
Barlow twins: Self-supervised learning via redundancy reduction,
J. Zbontar, L. Jing, I. Misra, Y . LeCun, and S. Deny, “Barlow twins: Self-supervised learning via redundancy reduction,” in Proc. 38th Inter. Conf. Machine Learning, (ICML) , vol. 139, 2021, pp. 12 310–12 320
work page 2021
-
[1]
Applications of Generative AI (GAI) for Mobile and Wireless Networking: A Survey,
T.-H. Vu, S. K. Jagatheesaperumal, M.-D. Nguyen, N. V . Huynh, S. Kim, and Q.-V . Pham, “Applications of Generative AI (GAI) for Mobile and Wireless Networking: A Survey,” IEEE Internet of Things Journ. , Aug. 2024
work page 2024
-
[2]
J. A. Cabrera, H. Boche, C. Deppe, R. F. Schaefer, C. Scheunert, and F. H. P. Fitzek, 6G and the post-Shannon theory . John Wiley & Sons, Ltd, 2022, pp. 271–294
work page 2022
-
[3]
AI empowered wireless communications: From bits to semantics,
Z. Qin, L. Liang, Z. Wang, S. Jin, X. Tao, W. Tong, and G. Y . Li, “AI empowered wireless communications: From bits to semantics,” Proc. IEEE, vol. 112, no. 7, pp. 621–652, 2024
work page 2024
-
[4]
Task-oriented communication design in cyber-physical systems: A survey on theory and applications,
A. Mostaani, T. X. Vu, S. K. Sharma, V .-D. Nguyen, Q. Liao, and S. Chatzinotas, “Task-oriented communication design in cyber-physical systems: A survey on theory and applications,” IEEE Access , vol. 10, pp. 133 842–133 868, 2022
work page 2022
-
[5]
Robust semantic communications against semantic noise,
Q. Hu, G. Zhang, Z. Qin, Y . Cai, G. Yu, and G. Y . Li, “Robust semantic communications against semantic noise,” in 2022 IEEE 96th Vehicular Technology Conference (VTC2022-Fall), 2022
work page 2022
-
[6]
Learning task-oriented communication for edge inference: An information bottleneck approach,
J. Shao, Y . Mao, and J. Zhang, “Learning task-oriented communication for edge inference: An information bottleneck approach,” IEEE Journal on Selected Areas in Communications, vol. 40, no. 1, pp. 197–211, 2021
work page 2021
-
[7]
Privacy-preserving task-oriented semantic communications against model inversion attacks,
Y . Wang, S. Guo, Y . Deng, H. Zhang, and Y . Fang, “Privacy-preserving task-oriented semantic communications against model inversion attacks,” IEEE Transactions on Wireless Communications , vol. 23, no. 8, pp. 10 150–10 165, 2024
work page 2024
Show all 62 references
-
[8]
Blockchain-based efficient and trustworthy aigc services in metaverse,
Y . Lin, Z. Gao, H. Du, D. Niyato, J. Kang, Z. Xiong, and Z. Zheng, “Blockchain-based efficient and trustworthy aigc services in metaverse,” IEEE Transactions on Services Computing , Mar. 2024
2024
-
[9]
A lite distributed semantic communication system for internet of things,
H. Xie and Z. Qin, “A lite distributed semantic communication system for internet of things,” IEEE Journal on Selected Areas in Communica- tions, vol. 39, no. 1, pp. 142–153, 2021
2021
-
[10]
Deep joint source- channel coding for wireless image transmission,
E. Bourtsoulatze, D. Burth Kurka, and D. G ¨und¨uz, “Deep joint source- channel coding for wireless image transmission,” IEEE Trans. Cogn. Commun. and Network. , vol. 5, no. 3, pp. 567–579, 2019
2019
-
[11]
Self-supervised learning: Generative or contrastive,
X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang, “Self-supervised learning: Generative or contrastive,”IEEE Transactions on Knowledge and Data Engineering, vol. 35, no. 1, pp. 857–876, 2023
2023
-
[12]
A mathematical theory of communication,
C. E. Shannon, “A mathematical theory of communication,” The Bell System Technical Journal, vol. 27, no. 3, pp. 379–423, 1948
1948
-
[13]
The Shannon sampling theorem—Its various extensions and applications: A tutorial review,
A. Jerri, “The Shannon sampling theorem—Its various extensions and applications: A tutorial review,” Proc. IEEE, vol. 65, no. 11, pp. 1565– 1596, 1977
1977
-
[14]
From semantic communication to semantic-aware networking: Model, architecture, and open problems,
G. Shi, Y . Xiao, Y . Li, and X. Xie, “From semantic communication to semantic-aware networking: Model, architecture, and open problems,” IEEE Communications Magazine , vol. 59, no. 8, pp. 44–50, 2021
2021
-
[15]
Internet of Things (IoT) communication protocols: Review,
S. Al-Sarawi, M. Anbar, K. Alieyan, and M. Alzubaidi, “Internet of Things (IoT) communication protocols: Review,” in Proc. 8th Inter. Confe. Infor. Tech. IEEE, 2017, pp. 685–690
2017
-
[16]
A survey on communication protocols and performance evaluations for internet of things,
C. Bayılmıs ¸, M. A. Ebleme, ¨Unal C ¸ avus ¸o˘glu, K. K ¨uc ¸¨uk, and A. Sevin, “A survey on communication protocols and performance evaluations for internet of things,” Digi. Commun. and Net. , vol. 8, no. 6, pp. 1094– 1104, 2022
2022
-
[17]
Research directions for the Internet of Things,
J. A. Stankovic, “Research directions for the Internet of Things,” IEEE Internet of Things J. , vol. 1, no. 1, pp. 3–9, 2014
2014
-
[18]
Deep learning enabled semantic communication systems,
H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,” IEEE Trans. Sign. Process. , Sep. 2021
2021
-
[19]
A Unified Multi- Task Semantic Communication System for Multimodal Data,
G. Zhang, Q. Hu, Z. Qin, Y . Cai, G. Yu, and X. Tao, “A Unified Multi- Task Semantic Communication System for Multimodal Data,” arXiv preprint arXiv:2209.07689, Aug. 2023
2023 arXiv
-
[20]
Task-Oriented Multi-User Semantic Communications,
H. Xie, Z. Qin, X. Tao, and K. B. Letaief, “Task-Oriented Multi-User Semantic Communications,” IEEE Jour. of Sel. Areas in Comm. , Mar. 2022
2022
-
[21]
Semantic communication with memory,
H. Xie, Z. Qin, and G. Y . Li, “Semantic communication with memory,” IEEE Jour. of Sel. Areas in Comm. , Jun. 2023
2023
-
[22]
Adaptable Semantic Com- pression and Resource Allocation for Task-Oriented Communications,
C. Liu, C. Guo, Y . Yang, and N. Jiang, “Adaptable Semantic Com- pression and Resource Allocation for Task-Oriented Communications,” IEEE Trans. Cogn. Comm. and Netw. , Jun. 2024. IEEE TRANSACTIONS ON MOBILE COMPUTING 15
2024
-
[23]
DeepJSCC-f: Deep Joint Source-Channel Coding of Images With Feedback,
D. B. Kurka and D. G ¨und¨uz, “DeepJSCC-f: Deep Joint Source-Channel Coding of Images With Feedback,” IEEE Jour. of Sel. Areas in Comm. , May 2020
2020
-
[24]
Deep Joint Source- Channel Coding for Wireless Image Transmission,
E. Bourtsoulatze, D. Burth Kurka, and D. G ¨und¨uz, “Deep Joint Source- Channel Coding for Wireless Image Transmission,” IEEE Trans. Cogn. Comm. and Netw. , Apr. 2019
2019
-
[25]
AdaSem: Adaptive Goal-Oriented Semantic Communications for End-to-End Camera Relocalization,
Q. Liao and T.-Y . Tung, “AdaSem: Adaptive Goal-Oriented Semantic Communications for End-to-End Camera Relocalization,” in INFOCOM, May 2024
2024
-
[26]
Contrastive learning-based semantic communications,
S. Tang, Q. Yang, L. Fan, X. Lei, A. Nallanathan, and G. K. Karagianni- dis, “Contrastive learning-based semantic communications,”IEEE Trans. on Comm., Oct. 2024
2024
-
[27]
DeepMA: End-to-end deep multiple access for wireless image transmission in semantic communication,
W. Zhang, K. Bai, S. Zeadally, H. Zhang, H. Shao, H. Ma, and V . C. M. Leung, “DeepMA: End-to-end deep multiple access for wireless image transmission in semantic communication,” IEEE Trans. Cogn. Comm. and Netw., Apr. 2024
2024
-
[28]
Generative joint source-channel coding for semantic image transmission,
E. Erdemir, T.-Y . Tung, P. L. Dragotti, and D. G¨und¨uz, “Generative joint source-channel coding for semantic image transmission,” IEEE Jour. of Sel. Areas in Comm. , Aug. 2023
2023
-
[29]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proceedings of the 38th International Conference on Machine L...
2021
-
[30]
Information constraints on auto-encoding variational bayes,
R. Lopez, J. Regier, M. I. Jordan, and N. Yosef, “Information constraints on auto-encoding variational bayes,” Advances in neural information processing systems, vol. 31, 2018
2018
-
[31]
Deep learning enabled semantic communication systems,
H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,” IEEE Trans. Signal Process., vol. 69, pp. 2663–2675, 2021
2021
-
[32]
Learning robust representations via multi-view information bottleneck,
M. Federici, A. Dutta, P. Forr ´e, N. Kushman, and Z. Akata, “Learning robust representations via multi-view information bottleneck,” in Inter- national Conference on Learning Representations , 2020
2020
-
[33]
To compress or not to compress—self- supervised learning and information theory: A review,
R. Shwartz Ziv and Y . LeCun, “To compress or not to compress—self- supervised learning and information theory: A review,” Entropy, vol. 26, no. 3, p. 252, 2024
2024
-
[34]
Deep learning and the information bot- tleneck principle,
N. Tishby and N. Zaslavsky, “Deep learning and the information bot- tleneck principle,” in 2015 IEEE Information Theory Workshop (ITW) , 2015, pp. 1–5
2015
-
[35]
Task-oriented communication with out-of-distribution detection: An information bottleneck framework,
H. Li, W. Yu, H. He, J. Shao, S. Song, J. Zhang, and K. B. Letaief, “Task-oriented communication with out-of-distribution detection: An information bottleneck framework,” in GLOBECOM 2023 - 2023 IEEE Global Communications Conference , Feb. 2024
2023
-
[36]
Discovering Invariant Rationales for Graph Neural Networks,
Y . Wu, X. Wang, A. Zhang, X. He, and T.-S. Chua, “Discovering Invariant Rationales for Graph Neural Networks,” in Int. Conf. Learn. Represent., May 2022
2022
-
[37]
Learning invariant representations and risks for semi-supervised domain adaptation,
B. Li, Y . Wang, S. Zhang, D. Li, K. Keutzer, T. Darrell, and H. Zhao, “Learning invariant representations and risks for semi-supervised domain adaptation,” in IEEE Conf. Comput. Vis. Pattern Recog. , Jun. 2021, pp. 1104–1113
2021
-
[38]
Rethinking infonce: How many negative samples do you need?
C. Wu, F. Wu, and Y . Huang, “Rethinking infonce: How many negative samples do you need?” in Proc. 31st Inter. Joint Conf. Artificial Intelligence, IJCAI-22, 7 2022, pp. 2509–2515
2022
-
[39]
VICReg: Variance-invariance- covariance regularization for self-supervised learning,
A. Bardes, J. Ponce, and Y . LeCun, “VICReg: Variance-invariance- covariance regularization for self-supervised learning,” in International Conference on Learning Representations , 2022
2022
-
[41]
Causality inspired representation learning for domain generalization,
F. Lv, J. Liang, S. Li, B. Zang, C. H. Liu, Z. Wang, and D. Liu, “Causality inspired representation learning for domain generalization,” in IEEE Conf. Comput. Vis. Pattern Recog. , June 2022, pp. 8046–8056
2022
-
[42]
Generalizable heterogeneous federated cross- correlation and instance similarity learning,
W. Huang et al. , “Generalizable heterogeneous federated cross- correlation and instance similarity learning,” IEEE Trans. Patt. Analysis and Machine Intelligence , vol. 46, no. 2, pp. 712–728, 2024
2024
-
[43]
A simple framework for contrastive learning of visual representations,
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in Proc. Inter. Conf. Machine Learning. PMLR, 2020, pp. 1597–1607
2020
-
[44]
Rethinking minimal sufficient representation in contrastive learning,
H. Wang, X. Guo, Z. Deng, and Y . Lu, “Rethinking minimal sufficient representation in contrastive learning,” in Proc. IEEE/CVF Conf. Com- put. Vision and Patt. Recog. (CVPR) , 2022, pp. 16 020–16 029
2022
-
[45]
Deep learning and the information bottleneck principle,
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in IEEE Infor. Theory Workshop (ITW), 2015, pp. 1–5
2015
-
[46]
Learning multiple layers of features from tiny images,
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” Department of Computer Science, University of Toronto, 2009
2009
-
[47]
EMNIST: Extending mnist to handwritten letters,
G. Cohen, S. Afshar, J. Tapson, and A. van Schaik, “EMNIST: Extending mnist to handwritten letters,” in 2017 International Joint Conference on Neural Networks (IJCNN) , 2017, pp. 2921–2926
2017
-
[48]
Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms,
H. Xiao, K. Rasul, and R. V ollgraf, “Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms,” arXiv preprint arXiv:1708.07747, 2017
2017 arXiv
-
[49]
An analysis of single-layer networks in unsupervised feature learning,
A. Coates, A. Ng, and H. Lee, “An analysis of single-layer networks in unsupervised feature learning,” in Proc. 14th Inter. Conf. Artificial Intelligence and Statistics, vol. 15. Fort Lauderdale, FL, USA: PMLR, 11–13 Apr 2011, pp. 215–223
2011
-
[50]
A visual vocabulary for flower clas- sification,
M.-E. Nilsback and A. Zisserman, “A visual vocabulary for flower clas- sification,” in IEEE Conf. Comput. Vision and Patt. Recog.(CVPR’06) , vol. 2, 2006, pp. 1447–1454
2006
-
[51]
The cityscapes dataset for semantic urban scene understanding,
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- nenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 3213– 3223
2016
-
[52]
Deeper, broader and artier domain generalization,
D. Li, Y . Yang, Y .-Z. Song, and T. M. Hospedales, “Deeper, broader and artier domain generalization,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5542–5550
2017
-
[54]
Gaussian error linear units (GELU),
D. Hendrycks and K. Gimpel, “Gaussian error linear units (GELU),” arXiv preprint arXiv:1606.08415 , 2016
2016 arXiv
-
[55]
Adam: A method for stochastic optimization,
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” CoRR, vol. abs/1412.6980, 2014
2014 arXiv
-
[56]
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in IEEE Inter. Conf. Computer Vision (ICCV) , 2015, pp. 1026–1034
2015
-
[57]
SGDR: Stochastic gradient descent with warm restarts,
I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in Inter. Conf. Learning Represent. , 2017
2017
-
[58]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778
2016
-
[59]
Contrastive learning-based semantic communications,
S. Tang, Q. Yang, L. Fan, X. Lei, A. Nallanathan, and G. K. Kara- giannidis, “Contrastive learning-based semantic communications,” IEEE Transactions on Communications, vol. 72, no. 10, pp. 6328–6343, 2024
2024
-
[60]
Deep learning enabled semantic communication systems,
H. Xie, Z. Qin, G. Y . Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,” IEEE Transactions on Signal Pro- cessing, vol. 69, pp. 2663–2675, 2021
2021
-
[61]
Segformer: Simple and efficient design for semantic segmentation with transformers,
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” Advances in neural information processing systems , vol. 34, pp. 12 077–12 090, 2021
2021
-
[62]
Refign: Align and refine for adaptation of semantic segmentation to adverse conditions,
D. Br ¨uggemann, C. Sakaridis, P. Truong, and L. Van Gool, “Refign: Align and refine for adaptation of semantic segmentation to adverse conditions,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 3174–3184
2023
-
[63]
Map-guided curriculum domain adaptation and uncertainty-aware evaluation for semantic nighttime im- age segmentation,
C. Sakaridis, D. Dai, and L. Van Gool, “Map-guided curriculum domain adaptation and uncertainty-aware evaluation for semantic nighttime im- age segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 6, pp. 3139–3153, 2020. Senura Hansaja Wa...
2020
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.